cdpfleet Docs GitHub Dashboard

Cases / #22 · 2026-10-06 · Easy

Check every link on a page: 73 links in seconds, from inside the browser

A link checker that stays the browser: collect the links, verify them with fetch() inside the page eight at a time, and compare with navigating to each one.

Chromium

Run on production on 2026-10-06: ✓ Node.js ✓ Python ✓ Java ✓ C# ✓ Go

The problem

Checking that every link on a page still works is the oldest crawl job there is, and the obvious browser version — open the page, click or navigate to each link, read the status — is also the slowest: every navigation loads a whole page with its stylesheets, scripts and images through your proxy. Doing the checks from your own HTTP client instead is fast, but then the checks come from a different IP, a different TLS stack and a different user agent than the browsing session (four ways to make a request). The middle path is to let the browser do the requests without navigating.

What we used, and why

WhatWhy
chromium, headless: "new"One session, one thread; nothing here is browser-specific.
page.$$eval("a[href]", …)Collects every link as an absolute URL, same-site only, de-duplicated, fragments dropped.
fetch(url, { cache: "no-store" }) inside page.evaluate, 8 URLs per callThe browser's network stack, proxy, cookies and headers — but no navigation, no rendering, no sub-resources. One round trip per wave of eight.
page.goto for a 10-link sampleThe baseline: what a user-like check costs per link.
page.on("request") counterShows what the navigations really fetched: not 10 pages but 288 requests.

How it works

  1. Open the front page and collect its same-site links (73 unique URLs).
  2. Check all 73 with fetch() inside the page, eight at a time; record status, redirects and failures.
  3. Navigate to the first 10 links one by one with page.goto, counting every request the browser makes.
  4. Print one row per method: links checked, OK / redirected / broken, seconds, seconds per link, requests made.

The code

The same program in five languages (also on GitHub, with the raw output). Set these environment variables first:

// npm install [email protected]
// env: CDPFLEET_API_KEY, PROXY_URL
import { chromium } from 'playwright';

const KEY = process.env.CDPFLEET_API_KEY;
const START = 'https://books.toscrape.com/'; // 70-odd same-site links on the front page
const WAVE = 8; // links checked at once

const res = await fetch('https://starter.cdpfleet.com/chromium/session', {
  method: 'POST',
  headers: { 'x-api-key': KEY, 'content-type': 'application/json' },
  body: JSON.stringify({ proxy: process.env.PROXY_URL, headless: 'new' }),
});
if (!res.ok) throw new Error(`launch ${res.status} ${await res.text()}`);
const { wsUrl } = await res.json();
const browser = await chromium.connect(wsUrl, { headers: { 'x-api-key': KEY } });

const summary = (method, results, seconds, extra = {}) => ({
  method,
  links: results.length,
  ok: results.filter((r) => r.status >= 200 && r.status < 300).length,
  redirected: results.filter((r) => r.redirected).length,
  broken: results.filter((r) => r.status === 0 || r.status >= 400).length,
  broken_urls: results.filter((r) => r.status === 0 || r.status >= 400).map((r) => r.url).slice(0, 5),
  seconds,
  seconds_per_link: Math.round((seconds / results.length) * 100) / 100,
  ...extra,
});

try {
  const page = await browser.newPage();
  await page.goto(START, { timeout: 60000 });
  // Every unique same-site link on the page, as absolute URLs.
  const links = await page.$$eval('a[href]', (as, origin) => [...new Set(as.map((a) => a.href))].filter((h) => h.startsWith(origin) && !h.includes('#')), new URL(START).origin);

  // Way 1: fetch() inside the page, WAVE links at a time — the browser's proxy, cookies
  // and TLS, no navigation, no rendering, no assets.
  const t1 = Date.now();
  const fetched = [];
  for (let i = 0; i < links.length; i += WAVE) {
    const wave = links.slice(i, i + WAVE);
    fetched.push(...await page.evaluate((urls) => Promise.all(urls.map(async (url) => {
      try { const r = await fetch(url, { cache: 'no-store' }); return { url, status: r.status, redirected: r.redirected }; } catch { return { url, status: 0, redirected: false }; }
    })), wave));
  }
  const inPage = summary(`fetch() in the page, ${WAVE} at a time`, fetched, (Date.now() - t1) / 1000);

  // Way 2: navigate to each link, like a user — full page loads with all their assets.
  // Only a sample: this is the slow way, and it is the same work for every link.
  const sample = links.slice(0, 10);
  let requests = 0;
  page.on('request', () => { requests++; });
  const t2 = Date.now();
  const navigated = [];
  for (const url of sample) {
    try { const r = await page.goto(url, { timeout: 60000 }); navigated.push({ url, status: r ? r.status() : 0, redirected: r ? r.url() !== url : false }); } catch { navigated.push({ url, status: 0, redirected: false }); }
  }
  const byNav = summary('page.goto each link (10-link sample)', navigated, (Date.now() - t2) / 1000, { requests_made: requests });

  console.log(JSON.stringify([{ ...inPage, requests_made: links.length }, byNav], null, 2));
} finally {
  await browser.close();
}

What we got

MethodLinksOKRedirectedBrokenSecondsPer link (s)Requests
fetch() in the page, 8 at a time7373007.6290.173
page.goto each link (10-link sample)10100022.2482.22288

From the Node.js run on 2026-10-06. IP addresses are replaced with placeholders (203.0.113.x); equal addresses stay equal. The other languages produced the same findings.

Raw output (Node.js)
[
  {
    "method": "fetch() in the page, 8 at a time",
    "links": 73,
    "ok": 73,
    "redirected": 0,
    "broken": 0,
    "broken_urls": [],
    "seconds": 7.629,
    "seconds_per_link": 0.1,
    "requests_made": 73
  },
  {
    "method": "page.goto each link (10-link sample)",
    "links": 10,
    "ok": 10,
    "redirected": 0,
    "broken": 0,
    "broken_urls": [],
    "seconds": 22.248,
    "seconds_per_link": 2.22,
    "requests_made": 288
  }
]

Takeaways