cdpfleet Docs GitHub Dashboard

Cases / #19 · 2026-10-05 · Easy

Fetch first, browser second: paying for JavaScript only when the page needs it

A plain HTTP GET through your proxy tells you in a second whether the data is in the HTML. Only the pages that render with JavaScript get a browser session — three pages, one decision rule, measured.

Chromium

Run on production on 2026-10-05: ✓ Node.js ✓ Python ✓ Java ✓ C# ✓ Go

The problem

A browser is the expensive way to download HTML. Many pages — catalogues, listings, documentation — ship their data in the HTML and need no JavaScript at all; others are empty shells filled in by scripts, and some need scrolling or clicks on top. If you send every URL to a browser you pay thread-seconds for pages curl could have fetched; if you send none you silently get empty results from the JavaScript ones. The question is how to decide per page, cheaply, through the same proxy, without maintaining two code paths.

What we used, and why

WhatWhy
playwright.request.newContext({ proxy })Playwright's HTTP client, run locally with your proxy — no session, no thread, same exit IP as the browser would have. Available in all five languages.
A browser-like User-Agent on the plain requestSo the server answers the plain GET the way it answers the browser; otherwise some sites serve a different page to a bare client.
Counting the item markup in the raw HTMLThe decision rule: if the HTML already contains the expected items (class="product_pod", class="quote"), there is nothing a browser would add.
chromium, headless: "new", launched lazilyOne session for all the pages that failed the rule; none at all if every page passed.
locator(…).first().waitFor() then count()In the browser, wait for the rendered items rather than for "load": JavaScript-rendered pages fire load before their content exists.

How it works

  1. GET each page through the proxy with the plain client; record status, size, how many items the HTML contains and the time.
  2. Pages whose HTML has fewer items than expected are marked as needing a browser.
  3. Launch one headless Chromium session only if that list is non-empty; open each such page, wait for the items to render, count them, record the time.
  4. Print one row per page: what the HTML had, what the browser had, which path it needed.

The code

The same program in five languages (also on GitHub, with the raw output). Set these environment variables first:

// npm install [email protected]
// env: CDPFLEET_API_KEY, PROXY_URL
import { chromium, request } from 'playwright';

const KEY = process.env.CDPFLEET_API_KEY;
const PROXY = new URL(process.env.PROXY_URL);

// Three pages, one question each: is the data in the HTML, or does it need JavaScript?
const TARGETS = [
  { url: 'https://books.toscrape.com/', item: 'product_pod', expected: 20 },
  { url: 'https://quotes.toscrape.com/js/', item: 'quote', expected: 10 },
  { url: 'https://quotes.toscrape.com/scroll', item: 'quote', expected: 10 },
];

// Step 1: a plain HTTP GET through the same proxy, no browser (Playwright's request API
// runs locally; the proxy keeps the exit IP identical to the browser's).
const http = await request.newContext({
  proxy: { server: `${PROXY.protocol}//${PROXY.host}`, username: decodeURIComponent(PROXY.username), password: decodeURIComponent(PROXY.password) },
  userAgent: 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/154.0.0.0 Safari/537.36',
});
const countInHtml = (html, cls) => (html.match(new RegExp(`class="[^"]*\\b${cls}\\b[^"]*"`, 'g')) || []).length;

const rows = [];
for (const t of TARGETS) {
  const t0 = Date.now();
  const res = await http.get(t.url, { timeout: 60000 });
  const html = await res.text();
  rows.push({ url: t.url, http_status: res.status(), html_kb: Math.round(html.length / 1024), items_in_html: countInHtml(html, t.item), fetch_seconds: (Date.now() - t0) / 1000 });
}
await http.dispose();

// Step 2: only the pages whose HTML didn't have the items get a browser.
const needsBrowser = rows.filter((r, i) => r.items_in_html < TARGETS[i].expected);
if (needsBrowser.length) {
  const launch = await fetch('https://starter.cdpfleet.com/chromium/session', {
    method: 'POST',
    headers: { 'x-api-key': KEY, 'content-type': 'application/json' },
    body: JSON.stringify({ proxy: process.env.PROXY_URL, headless: 'new' }),
  });
  if (!launch.ok) throw new Error(`launch ${launch.status} ${await launch.text()}`);
  const { wsUrl } = await launch.json();
  const browser = await chromium.connect(wsUrl, { headers: { 'x-api-key': KEY } });
  try {
    for (const r of needsBrowser) {
      const t = TARGETS.find((x) => x.url === r.url);
      const page = await browser.newPage();
      const t0 = Date.now();
      await page.goto(r.url, { timeout: 60000 });
      await page.locator(`.${t.item}`).first().waitFor({ timeout: 60000 });
      r.items_in_browser = await page.locator(`.${t.item}`).count();
      r.browser_seconds = (Date.now() - t0) / 1000;
      await page.close();
    }
  } finally {
    await browser.close();
  }
}
for (const r of rows) {
  r.needs_browser = r.items_in_browser !== undefined;
  r.items_in_browser ??= null;
  r.browser_seconds ??= null;
}
console.log(JSON.stringify(rows, null, 2));

What we got

PageStatusHTML (KB)Items in HTMLItems in browserNeeded a browserFetch (s)Browser (s)
https://books.toscrape.com/2005020—no3.55—
https://quotes.toscrape.com/js/2006010yes1.22311.043
https://quotes.toscrape.com/scroll2003010yes2.6056.12

From the Node.js run on 2026-10-05. IP addresses are replaced with placeholders (203.0.113.x); equal addresses stay equal. The other languages produced the same findings.

Raw output (Node.js)
[
  {
    "url": "https://books.toscrape.com/",
    "http_status": 200,
    "html_kb": 50,
    "items_in_html": 20,
    "fetch_seconds": 3.55,
    "needs_browser": false,
    "items_in_browser": null,
    "browser_seconds": null
  },
  {
    "url": "https://quotes.toscrape.com/js/",
    "http_status": 200,
    "html_kb": 6,
    "items_in_html": 0,
    "fetch_seconds": 1.223,
    "items_in_browser": 10,
    "browser_seconds": 11.043,
    "needs_browser": true
  },
  {
    "url": "https://quotes.toscrape.com/scroll",
    "http_status": 200,
    "html_kb": 3,
    "items_in_html": 0,
    "fetch_seconds": 2.605,
    "items_in_browser": 10,
    "browser_seconds": 6.12,
    "needs_browser": true
  }
]

Takeaways