cdpfleet Docs GitHub Dashboard

Cases / #23 · 2026-10-06 · Easy

Download every image on a page without downloading it twice

Two ways to get a page's images as bytes on your machine: keep the responses the page already loaded (zero extra requests), or fetch them again from inside the page. Both through the browser, both verified as real JPEGs.

Chromium

Run on production on 2026-10-06: ✓ Node.js ✓ Python ✓ Java ✓ C# ✓ Go

The problem

Saving the images from a page sounds like a one-liner until you think about where the bytes should come from. Download the URLs with your own HTTP client and you hit the CDN from another IP with another fingerprint and no cookies — hotlink protection and signed URLs are built for exactly that. Fetch them again in the browser and you pay for every image twice through your proxy. But the page already downloaded them: can you just keep those bytes?

What we used, and why

WhatWhy
chromium, headless: "new"One session, one thread.
page.on("response") with request().resourceType() === "image" and response.body()Keeps the bytes of every image the page loads, as it loads them — no second request.
waitUntil: "networkidle"So all cover images have arrived before we count them.
$$eval("article.product_pod img", …) reading currentSrcThe list of images we actually want, matched against what was captured.
fetch(url) → arrayBuffer() → base64 inside page.evaluateThe re-fetch alternative: same cookies, proxy and headers as the page, bytes handed to the client as text.
The JPEG magic bytes FF D8 FFProof the bytes are the image, not an error page.

How it works

  1. Open the page with a response listener that stores the body of every successful image response.
  2. List the 20 cover images on the page and match them against the captured responses.
  3. Fetch the same 20 images again from inside the page and decode them on the client.
  4. For each method: count images, validate JPEG headers, sum the bytes, count extra requests, time it.

The code

The same program in five languages (also on GitHub, with the raw output). Set these environment variables first:

// npm install [email protected]
// env: CDPFLEET_API_KEY, PROXY_URL
import { chromium } from 'playwright';

const KEY = process.env.CDPFLEET_API_KEY;
const PAGE = 'https://books.toscrape.com/'; // 20 cover images

const res = await fetch('https://starter.cdpfleet.com/chromium/session', {
  method: 'POST',
  headers: { 'x-api-key': KEY, 'content-type': 'application/json' },
  body: JSON.stringify({ proxy: process.env.PROXY_URL, headless: 'new' }),
});
if (!res.ok) throw new Error(`launch ${res.status} ${await res.text()}`);
const { wsUrl } = await res.json();
const browser = await chromium.connect(wsUrl, { headers: { 'x-api-key': KEY } });

const isJpeg = (buf) => buf.length > 3 && buf[0] === 0xff && buf[1] === 0xd8 && buf[2] === 0xff;
const summary = (method, files, seconds, extra_requests) => ({
  method, images: files.length, valid_jpeg: files.filter((f) => isJpeg(f.bytes)).length,
  total_kb: Math.round(files.reduce((n, f) => n + f.bytes.length, 0) / 1024), extra_requests, seconds,
});

try {
  // Way 1: keep the bytes the page downloads anyway — zero extra requests.
  const page = await browser.newPage();
  const captured = [];
  page.on('response', async (r) => {
    if (r.request().resourceType() === 'image' && r.ok()) {
      try { captured.push({ url: r.url(), bytes: await r.body() }); } catch { /* body gone (cache) */ }
    }
  });
  const t1 = Date.now();
  await page.goto(PAGE, { timeout: 60000, waitUntil: 'networkidle' });
  const covers = await page.$$eval('article.product_pod img', (imgs) => imgs.map((i) => i.currentSrc || i.src));
  const fromLoad = captured.filter((c) => covers.includes(c.url));
  const way1 = summary('capture responses while the page loads', fromLoad, (Date.now() - t1) / 1000, 0);

  // Way 2: fetch each image again from inside the page (same cookies, proxy and headers as
  // the page itself) and hand the bytes to the client as base64.
  const t2 = Date.now();
  const again = await page.evaluate((urls) => Promise.all(urls.map(async (url) => {
    const buf = await (await fetch(url, { cache: 'no-store' })).arrayBuffer();
    let s = ''; const b = new Uint8Array(buf); for (let i = 0; i < b.length; i++) s += String.fromCharCode(b[i]);
    return { url, b64: btoa(s) };
  })), covers);
  const refetched = again.map((a) => ({ url: a.url, bytes: Buffer.from(a.b64, 'base64') }));
  const way2 = summary('fetch() each image again in the page', refetched, (Date.now() - t2) / 1000, covers.length);

  console.log(JSON.stringify([way1, way2], null, 2));
} finally {
  await browser.close();
}

What we got

MethodImagesValid JPEGTotal (KB)Extra requestsSeconds
capture responses while the page loads2020173022.159
fetch() each image again in the page20201732010.317

From the Node.js run on 2026-10-06. IP addresses are replaced with placeholders (203.0.113.x); equal addresses stay equal. The other languages produced the same findings.

Raw output (Node.js)
[
  {
    "method": "capture responses while the page loads",
    "images": 20,
    "valid_jpeg": 20,
    "total_kb": 173,
    "extra_requests": 0,
    "seconds": 22.159
  },
  {
    "method": "fetch() each image again in the page",
    "images": 20,
    "valid_jpeg": 20,
    "total_kb": 173,
    "extra_requests": 20,
    "seconds": 10.317
  }
]

Takeaways