Cases / #23 · 2026-10-06 · Easy
Download every image on a page without downloading it twice
Two ways to get a page's images as bytes on your machine: keep the responses the page already loaded (zero extra requests), or fetch them again from inside the page. Both through the browser, both verified as real JPEGs.
Run on production on 2026-10-06: ✓ Node.js ✓ Python ✓ Java ✓ C# ✓ Go
The problem
Saving the images from a page sounds like a one-liner until you think about where the bytes should come from. Download the URLs with your own HTTP client and you hit the CDN from another IP with another fingerprint and no cookies — hotlink protection and signed URLs are built for exactly that. Fetch them again in the browser and you pay for every image twice through your proxy. But the page already downloaded them: can you just keep those bytes?
What we used, and why
| What | Why |
|---|---|
chromium, headless: "new" | One session, one thread. |
page.on("response") with request().resourceType() === "image" and response.body() | Keeps the bytes of every image the page loads, as it loads them — no second request. |
waitUntil: "networkidle" | So all cover images have arrived before we count them. |
$$eval("article.product_pod img", …) reading currentSrc | The list of images we actually want, matched against what was captured. |
fetch(url) → arrayBuffer() → base64 inside page.evaluate | The re-fetch alternative: same cookies, proxy and headers as the page, bytes handed to the client as text. |
The JPEG magic bytes FF D8 FF | Proof the bytes are the image, not an error page. |
How it works
- Open the page with a response listener that stores the body of every successful image response.
- List the 20 cover images on the page and match them against the captured responses.
- Fetch the same 20 images again from inside the page and decode them on the client.
- For each method: count images, validate JPEG headers, sum the bytes, count extra requests, time it.
The code
The same program in five languages (also on GitHub, with the raw output). Set these environment variables first:
CDPFLEET_API_KEY— your API key (dashboard → API keys)PROXY_URL— your proxy, e.g.http://user:[email protected]:8000
// npm install [email protected]
// env: CDPFLEET_API_KEY, PROXY_URL
import { chromium } from 'playwright';
const KEY = process.env.CDPFLEET_API_KEY;
const PAGE = 'https://books.toscrape.com/'; // 20 cover images
const res = await fetch('https://starter.cdpfleet.com/chromium/session', {
method: 'POST',
headers: { 'x-api-key': KEY, 'content-type': 'application/json' },
body: JSON.stringify({ proxy: process.env.PROXY_URL, headless: 'new' }),
});
if (!res.ok) throw new Error(`launch ${res.status} ${await res.text()}`);
const { wsUrl } = await res.json();
const browser = await chromium.connect(wsUrl, { headers: { 'x-api-key': KEY } });
const isJpeg = (buf) => buf.length > 3 && buf[0] === 0xff && buf[1] === 0xd8 && buf[2] === 0xff;
const summary = (method, files, seconds, extra_requests) => ({
method, images: files.length, valid_jpeg: files.filter((f) => isJpeg(f.bytes)).length,
total_kb: Math.round(files.reduce((n, f) => n + f.bytes.length, 0) / 1024), extra_requests, seconds,
});
try {
// Way 1: keep the bytes the page downloads anyway — zero extra requests.
const page = await browser.newPage();
const captured = [];
page.on('response', async (r) => {
if (r.request().resourceType() === 'image' && r.ok()) {
try { captured.push({ url: r.url(), bytes: await r.body() }); } catch { /* body gone (cache) */ }
}
});
const t1 = Date.now();
await page.goto(PAGE, { timeout: 60000, waitUntil: 'networkidle' });
const covers = await page.$$eval('article.product_pod img', (imgs) => imgs.map((i) => i.currentSrc || i.src));
const fromLoad = captured.filter((c) => covers.includes(c.url));
const way1 = summary('capture responses while the page loads', fromLoad, (Date.now() - t1) / 1000, 0);
// Way 2: fetch each image again from inside the page (same cookies, proxy and headers as
// the page itself) and hand the bytes to the client as base64.
const t2 = Date.now();
const again = await page.evaluate((urls) => Promise.all(urls.map(async (url) => {
const buf = await (await fetch(url, { cache: 'no-store' })).arrayBuffer();
let s = ''; const b = new Uint8Array(buf); for (let i = 0; i < b.length; i++) s += String.fromCharCode(b[i]);
return { url, b64: btoa(s) };
})), covers);
const refetched = again.map((a) => ({ url: a.url, bytes: Buffer.from(a.b64, 'base64') }));
const way2 = summary('fetch() each image again in the page', refetched, (Date.now() - t2) / 1000, covers.length);
console.log(JSON.stringify([way1, way2], null, 2));
} finally {
await browser.close();
}
# pip install playwright==1.60.0 requests
# env: CDPFLEET_API_KEY, PROXY_URL
import base64
import json
import os
import time
import requests
from playwright.sync_api import sync_playwright
KEY = os.environ["CDPFLEET_API_KEY"]
PAGE = "https://books.toscrape.com/" # 20 cover images
res = requests.post("https://starter.cdpfleet.com/chromium/session", headers={"x-api-key": KEY},
json={"proxy": os.environ["PROXY_URL"], "headless": "new"}, timeout=60)
res.raise_for_status()
def is_jpeg(b):
return len(b) > 3 and b[0] == 0xFF and b[1] == 0xD8 and b[2] == 0xFF
def summary(method, files, seconds, extra_requests):
return {
"method": method, "images": len(files), "valid_jpeg": sum(1 for f in files if is_jpeg(f["bytes"])),
"total_kb": int(sum(len(f["bytes"]) for f in files) / 1024 + 0.5), "extra_requests": extra_requests, "seconds": seconds,
}
# Fetch each image again from inside the page and hand the bytes over as base64.
REFETCH = """urls => Promise.all(urls.map(async (url) => {
const buf = await (await fetch(url, { cache: 'no-store' })).arrayBuffer();
let s = ''; const b = new Uint8Array(buf); for (let i = 0; i < b.length; i++) s += String.fromCharCode(b[i]);
return { url, b64: btoa(s) };
}))"""
with sync_playwright() as p:
browser = p.chromium.connect(res.json()["wsUrl"], headers={"x-api-key": KEY})
try:
# Way 1: keep the bytes the page downloads anyway — zero extra requests. In the sync
# API the handler only remembers the response; bodies are read after navigation.
page = browser.new_page()
image_responses = []
page.on("response", lambda r: image_responses.append(r) if r.request.resource_type == "image" and r.ok else None)
t1 = time.time()
page.goto(PAGE, timeout=60000, wait_until="networkidle")
covers = page.eval_on_selector_all("article.product_pod img", "imgs => imgs.map(i => i.currentSrc || i.src)")
captured = []
for r in image_responses:
try:
captured.append({"url": r.url, "bytes": r.body()})
except Exception:
pass # body gone (cache)
from_load = [c for c in captured if c["url"] in covers]
way1 = summary("capture responses while the page loads", from_load, round(time.time() - t1, 3), 0)
# Way 2: fetch each image again from inside the page (same cookies, proxy and headers).
t2 = time.time()
again = page.evaluate(REFETCH, covers)
refetched = [{"url": a["url"], "bytes": base64.b64decode(a["b64"])} for a in again]
way2 = summary("fetch() each image again in the page", refetched, round(time.time() - t2, 3), len(covers))
print(json.dumps([way1, way2], indent=2))
finally:
browser.close()
// Maven: com.microsoft.playwright:playwright:1.60.0, com.google.code.gson:gson:2.11.0
// Run with PLAYWRIGHT_SKIP_BROWSER_DOWNLOAD=1. env: CDPFLEET_API_KEY, PROXY_URL
import com.google.gson.*;
import com.microsoft.playwright.*;
import com.microsoft.playwright.options.WaitUntilState;
import java.net.URI;
import java.net.http.*;
import java.util.*;
public class Main {
static final String KEY = System.getenv("CDPFLEET_API_KEY");
static final String PAGE = "https://books.toscrape.com/"; // 20 cover images
// Fetch each image again from inside the page and hand the bytes over as base64.
static final String REFETCH = """
urls => Promise.all(urls.map(async (url) => {
const buf = await (await fetch(url, { cache: 'no-store' })).arrayBuffer();
let s = ''; const b = new Uint8Array(buf); for (let i = 0; i < b.length; i++) s += String.fromCharCode(b[i]);
return { url, b64: btoa(s) };
}))""";
record Img(String url, byte[] bytes) {}
static boolean isJpeg(byte[] b) {
return b.length > 3 && (b[0] & 0xff) == 0xff && (b[1] & 0xff) == 0xd8 && (b[2] & 0xff) == 0xff;
}
static JsonObject summary(String method, List<Img> files, double seconds, int extraRequests) {
long total = 0;
int valid = 0;
for (Img f : files) { total += f.bytes().length; if (isJpeg(f.bytes())) valid++; }
JsonObject o = new JsonObject();
o.addProperty("method", method);
o.addProperty("images", files.size());
o.addProperty("valid_jpeg", valid);
o.addProperty("total_kb", Math.round(total / 1024.0));
o.addProperty("extra_requests", extraRequests);
o.addProperty("seconds", seconds);
return o;
}
public static void main(String[] args) throws Exception {
HttpClient http = HttpClient.newHttpClient();
String body = "{\"proxy\": " + new Gson().toJson(System.getenv("PROXY_URL")) + ", \"headless\": \"new\"}";
HttpResponse<String> res = http.send(HttpRequest.newBuilder(URI.create("https://starter.cdpfleet.com/chromium/session"))
.header("x-api-key", KEY).header("content-type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString(body)).build(), HttpResponse.BodyHandlers.ofString());
if (res.statusCode() != 200) throw new RuntimeException("launch: " + res.statusCode() + " " + res.body());
String wsUrl = JsonParser.parseString(res.body()).getAsJsonObject().get("wsUrl").getAsString();
try (Playwright playwright = Playwright.create()) {
Browser browser = playwright.chromium().connect(wsUrl, new BrowserType.ConnectOptions().setHeaders(Map.of("x-api-key", KEY)));
try {
// Way 1: keep the bytes the page downloads anyway — zero extra requests. The listener
// only remembers the response; bodies are read after navigation.
Page page = browser.newPage();
List<Response> imageResponses = Collections.synchronizedList(new ArrayList<>());
page.onResponse(r -> { if ("image".equals(r.request().resourceType()) && r.ok()) imageResponses.add(r); });
long t1 = System.currentTimeMillis();
page.navigate(PAGE, new Page.NavigateOptions().setTimeout(60000).setWaitUntil(WaitUntilState.NETWORKIDLE));
@SuppressWarnings("unchecked")
List<String> covers = (List<String>) page.evalOnSelectorAll("article.product_pod img", "imgs => imgs.map(i => i.currentSrc || i.src)");
List<Img> fromLoad = new ArrayList<>();
for (Response r : new ArrayList<>(imageResponses)) {
try { if (covers.contains(r.url())) fromLoad.add(new Img(r.url(), r.body())); } catch (PlaywrightException e) { /* body gone (cache) */ }
}
JsonObject way1 = summary("capture responses while the page loads", fromLoad, (System.currentTimeMillis() - t1) / 1000.0, 0);
// Way 2: fetch each image again from inside the page (same cookies, proxy and headers).
long t2 = System.currentTimeMillis();
@SuppressWarnings("unchecked")
List<Map<String, Object>> again = (List<Map<String, Object>>) page.evaluate(REFETCH, covers);
List<Img> refetched = new ArrayList<>();
for (Map<String, Object> a : again) refetched.add(new Img((String) a.get("url"), Base64.getDecoder().decode((String) a.get("b64"))));
JsonObject way2 = summary("fetch() each image again in the page", refetched, (System.currentTimeMillis() - t2) / 1000.0, covers.size());
JsonArray out = new JsonArray();
out.add(way1);
out.add(way2);
System.out.println(new GsonBuilder().setPrettyPrinting().disableHtmlEscaping().serializeNulls().create().toJson(out));
} finally {
browser.close();
}
}
}
}
// dotnet add package Microsoft.Playwright --version 1.60.0
// env: CDPFLEET_API_KEY, PROXY_URL
using System.Diagnostics;
using System.Net.Http.Json;
using System.Text.Encodings.Web;
using System.Text.Json;
using System.Text.Json.Nodes;
using Microsoft.Playwright;
var key = Environment.GetEnvironmentVariable("CDPFLEET_API_KEY")!;
const string PageUrl = "https://books.toscrape.com/"; // 20 cover images
// Fetch each image again from inside the page and hand the bytes over as base64.
const string Refetch = """
urls => Promise.all(urls.map(async (url) => {
const buf = await (await fetch(url, { cache: 'no-store' })).arrayBuffer();
let s = ''; const b = new Uint8Array(buf); for (let i = 0; i < b.length; i++) s += String.fromCharCode(b[i]);
return { url, b64: btoa(s) };
}))
""";
static bool IsJpeg(byte[] b) => b.Length > 3 && b[0] == 0xFF && b[1] == 0xD8 && b[2] == 0xFF;
static JsonObject Summary(string method, List<(string Url, byte[] Bytes)> files, double seconds, int extraRequests) => new()
{
["method"] = method,
["images"] = files.Count,
["valid_jpeg"] = files.Count(f => IsJpeg(f.Bytes)),
["total_kb"] = (int)Math.Floor(files.Sum(f => (long)f.Bytes.Length) / 1024.0 + 0.5),
["extra_requests"] = extraRequests,
["seconds"] = seconds,
};
using var http = new HttpClient();
http.DefaultRequestHeaders.Add("x-api-key", key);
var res = await http.PostAsJsonAsync("https://starter.cdpfleet.com/chromium/session",
new { proxy = Environment.GetEnvironmentVariable("PROXY_URL"), headless = "new" });
if (!res.IsSuccessStatusCode) throw new Exception($"launch: {(int)res.StatusCode} {await res.Content.ReadAsStringAsync()}");
var wsUrl = (await res.Content.ReadFromJsonAsync<JsonElement>()).GetProperty("wsUrl").GetString()!;
using var playwright = await Playwright.CreateAsync();
var browser = await playwright.Chromium.ConnectAsync(wsUrl, new() { Headers = new Dictionary<string, string> { ["x-api-key"] = key } });
try
{
// Way 1: keep the bytes the page downloads anyway — zero extra requests. The handler
// only remembers the response; bodies are read after navigation.
var page = await browser.NewPageAsync();
var imageResponses = new List<IResponse>();
page.Response += (_, r) => { if (r.Request.ResourceType == "image" && r.Ok) lock (imageResponses) imageResponses.Add(r); };
var sw = Stopwatch.StartNew();
await page.GotoAsync(PageUrl, new() { Timeout = 60000, WaitUntil = WaitUntilState.NetworkIdle });
var covers = (await page.EvalOnSelectorAllAsync<string[]>("article.product_pod img", "imgs => imgs.map(i => i.currentSrc || i.src)")).ToList();
var fromLoad = new List<(string, byte[])>();
IResponse[] snapshot; lock (imageResponses) snapshot = imageResponses.ToArray();
foreach (var r in snapshot)
{
if (!covers.Contains(r.Url)) continue;
try { fromLoad.Add((r.Url, await r.BodyAsync())); } catch (PlaywrightException) { /* body gone (cache) */ }
}
var way1 = Summary("capture responses while the page loads", fromLoad, sw.ElapsedMilliseconds / 1000.0, 0);
// Way 2: fetch each image again from inside the page (same cookies, proxy and headers).
sw.Restart();
var again = await page.EvaluateAsync<JsonElement>(Refetch, covers);
var refetched = again.EnumerateArray().Select(a => (a.GetProperty("url").GetString()!, Convert.FromBase64String(a.GetProperty("b64").GetString()!))).ToList();
var way2 = Summary("fetch() each image again in the page", refetched, sw.ElapsedMilliseconds / 1000.0, covers.Count);
var output = new JsonArray(way1, way2);
Console.WriteLine(output.ToJsonString(new JsonSerializerOptions { WriteIndented = true, Encoder = JavaScriptEncoder.UnsafeRelaxedJsonEscaping }));
}
finally
{
await browser.CloseAsync();
}
// go get github.com/playwright-community/[email protected]
// Driver: build playwright-core 1.60.0 from npm and set PLAYWRIGHT_DRIVER_PATH (see /docs/quickstart).
// env: CDPFLEET_API_KEY, PROXY_URL
package main
import (
"bytes"
"encoding/base64"
"encoding/json"
"fmt"
"io"
"log"
"math"
"net/http"
"os"
"sync"
"time"
"github.com/playwright-community/playwright-go"
)
var key = os.Getenv("CDPFLEET_API_KEY")
const pageURL = "https://books.toscrape.com/" // 20 cover images
// Fetch each image again from inside the page and hand the bytes over as base64.
const refetch = `urls => Promise.all(urls.map(async (url) => {
const buf = await (await fetch(url, { cache: 'no-store' })).arrayBuffer();
let s = ''; const b = new Uint8Array(buf); for (let i = 0; i < b.length; i++) s += String.fromCharCode(b[i]);
return { url, b64: btoa(s) };
}))`
type img struct {
url string
bytes []byte
}
type row struct {
Method string `json:"method"`
Images int `json:"images"`
ValidJpeg int `json:"valid_jpeg"`
TotalKB int `json:"total_kb"`
ExtraRequests int `json:"extra_requests"`
Seconds float64 `json:"seconds"`
}
func launch(name string, options map[string]any) (map[string]any, error) {
body, _ := json.Marshal(options)
req, _ := http.NewRequest("POST", "https://starter.cdpfleet.com/"+name+"/session", bytes.NewReader(body))
req.Header.Set("x-api-key", key)
req.Header.Set("content-type", "application/json")
res, err := http.DefaultClient.Do(req)
if err != nil {
return nil, err
}
defer res.Body.Close()
if res.StatusCode != http.StatusOK {
msg, _ := io.ReadAll(res.Body)
return nil, fmt.Errorf("launch %s: %s %s", name, res.Status, msg)
}
var session map[string]any
return session, json.NewDecoder(res.Body).Decode(&session)
}
func must[T any](v T, err error) T {
if err != nil {
log.Fatal(err)
}
return v
}
func isJpeg(b []byte) bool { return len(b) > 3 && b[0] == 0xff && b[1] == 0xd8 && b[2] == 0xff }
func summary(method string, files []img, seconds float64, extra int) row {
total, valid := 0, 0
for _, f := range files {
total += len(f.bytes)
if isJpeg(f.bytes) {
valid++
}
}
return row{Method: method, Images: len(files), ValidJpeg: valid, TotalKB: int(math.Floor(float64(total)/1024 + 0.5)), ExtraRequests: extra, Seconds: seconds}
}
func main() {
session := must(launch("chromium", map[string]any{"proxy": os.Getenv("PROXY_URL"), "headless": "new"}))
pw := must(playwright.Run(&playwright.RunOptions{SkipInstallBrowsers: true}))
defer pw.Stop()
browser := must(pw.Chromium.Connect(session["wsUrl"].(string), playwright.BrowserTypeConnectOptions{Headers: map[string]string{"x-api-key": key}}))
defer browser.Close()
// Way 1: keep the bytes the page downloads anyway — zero extra requests. The listener
// only remembers the response; bodies are read after navigation.
page := must(browser.NewPage())
var mu sync.Mutex
var imageResponses []playwright.Response
page.OnResponse(func(r playwright.Response) {
if r.Request().ResourceType() == "image" && r.Ok() {
mu.Lock()
imageResponses = append(imageResponses, r)
mu.Unlock()
}
})
t1 := time.Now()
must(page.Goto(pageURL, playwright.PageGotoOptions{Timeout: playwright.Float(60000), WaitUntil: playwright.WaitUntilStateNetworkidle}))
raw := must(page.Locator("article.product_pod img").EvaluateAll("imgs => imgs.map(i => i.currentSrc || i.src)"))
covers := map[string]bool{}
var coverList []string
for _, u := range raw.([]any) {
covers[u.(string)] = true
coverList = append(coverList, u.(string))
}
var fromLoad []img
mu.Lock()
snapshot := append([]playwright.Response(nil), imageResponses...)
mu.Unlock()
for _, r := range snapshot {
if !covers[r.URL()] {
continue
}
if b, err := r.Body(); err == nil { // body gone (cache) → skip
fromLoad = append(fromLoad, img{r.URL(), b})
}
}
way1 := summary("capture responses while the page loads", fromLoad, float64(time.Since(t1).Milliseconds())/1000, 0)
// Way 2: fetch each image again from inside the page (same cookies, proxy and headers).
t2 := time.Now()
again := must(page.Evaluate(refetch, coverList))
var refetched []img
for _, a := range again.([]any) {
m := a.(map[string]any)
refetched = append(refetched, img{m["url"].(string), must(base64.StdEncoding.DecodeString(m["b64"].(string)))})
}
way2 := summary("fetch() each image again in the page", refetched, float64(time.Since(t2).Milliseconds())/1000, len(coverList))
out, _ := json.MarshalIndent([]row{way1, way2}, "", " ")
fmt.Println(string(out))
}
What we got
| Method | Images | Valid JPEG | Total (KB) | Extra requests | Seconds |
|---|---|---|---|---|---|
| capture responses while the page loads | 20 | 20 | 173 | 0 | 22.159 |
| fetch() each image again in the page | 20 | 20 | 173 | 20 | 10.317 |
From the Node.js run on 2026-10-06. IP addresses are replaced with placeholders (203.0.113.x); equal addresses stay equal. The other languages produced the same findings.
Raw output (Node.js)
[
{
"method": "capture responses while the page loads",
"images": 20,
"valid_jpeg": 20,
"total_kb": 173,
"extra_requests": 0,
"seconds": 22.159
},
{
"method": "fetch() each image again in the page",
"images": 20,
"valid_jpeg": 20,
"total_kb": 173,
"extra_requests": 20,
"seconds": 10.317
}
]Takeaways
- Capturing responses gives you all 20 images with zero extra requests — the bytes were already on their way to the renderer;
response.body()streams them to your client over the fleet connection. - Both methods produce the identical 173 KB of valid JPEGs, so re-fetching buys nothing except a second trip through the proxy for every file.
- Capture while loading, fetch for the rest: capture is free for everything the page displays; use the in-page fetch for images the page only links to (full-size originals behind thumbnails) — still as the browser, with its cookies and proxy.
- The timing difference here is page load, not method: the capture number includes waiting for
networkidlethrough the proxy; the re-fetch ran on a warm page. Count requests and bytes, which are the real cost. - Watch for cached responses:
response.body()can fail for responses served from cache; the capture code ignores those, which is why the list is matched against the page's images.