cdpfleet Docs GitHub Dashboard

Cases / #17 · 2026-10-04 · Easy

Infinite scroll, two ways: scroll like a user, or call the JSON the page calls

Collecting everything an endlessly scrolling page loads: by scrolling and watching the DOM, and by catching the page's own API request and paging through it from inside the browser.

Chromium

Run on production on 2026-10-04: ✓ Node.js ✓ Python ✓ Java ✓ C# ✓ Go

The problem

Infinite-scroll pages have no "next" link and no end: content appears as you scroll, and the only way to know you've got everything is that scrolling stops producing more. The naive script scrolls, waits, counts, repeats — and has to guess how long to wait. But the page gets its data from somewhere: an XHR to a JSON endpoint, usually paginated, usually unauthenticated. Which approach is faster, which is more reliable, and how do you call that endpoint without losing the browser's proxy, cookies and TLS fingerprint?

What we used, and why

WhatWhy
chromium, headless: "new"Nothing here depends on the browser; the cheapest session (1 thread) does.
page.on("request")Counts every request each method causes — page assets included — so the two can be compared.
window.scrollTo(0, document.body.scrollHeight) in a loopThe user-like way: scroll, wait, count .quote elements, stop after three scrolls that added nothing.
page.waitForResponse(xhr)Catches the first JSON request the page makes for its own content, which gives us the endpoint, its parameters and page 1 of the data.
page.evaluate(urls => Promise.all(urls.map(fetch)))Calls the endpoint from inside the page: the browser's proxy, cookies, headers and TLS fingerprint, four pages per round trip.
A fresh tab per methodSo request counts and timings don't mix.

How it works

  1. Open the page and wait for the first quotes to render; record the load time separately from the collection time.
  2. Scroll: scroll to the bottom, wait 600 ms, count the quotes; stop when three scrolls in a row add nothing; read the quotes from the DOM.
  3. API: in a new tab, wait for the page's first XHR to /api/…, take page 1 from its body, then fetch pages 2–5, 6–9, 10–13 in waves of four with fetch inside the page until a wave reports no next page.
  4. Print, per method: quotes and distinct authors found, scrolls or API pages, requests made, load and collection seconds.

The code

The same program in five languages (also on GitHub, with the raw output). Set these environment variables first:

// npm install [email protected]
// env: CDPFLEET_API_KEY, PROXY_URL
import { chromium } from 'playwright';

const KEY = process.env.CDPFLEET_API_KEY;
const START = 'https://quotes.toscrape.com/scroll'; // loads 10 quotes per screen, 100 in all

// Way 1: behave like a user — scroll to the bottom until nothing more appears, then read the DOM.
async function byScrolling(page) {
  const t = Date.now();
  let requests = 0;
  page.on('request', () => { requests++; });
  await page.goto(START, { timeout: 60000 });
  await page.locator('.quote').first().waitFor({ timeout: 60000 });
  const loaded = Date.now();
  let count = 0;
  let scrolls = 0;
  let stale = 0;
  while (stale < 3) {
    await page.evaluate(() => window.scrollTo(0, document.body.scrollHeight));
    scrolls++;
    await page.waitForTimeout(600);
    const now = await page.locator('.quote').count();
    stale = now > count ? 0 : stale + 1;
    count = now;
  }
  const quotes = await page.$$eval('.quote', (els) => els.map((e) => ({ text: e.querySelector('.text').textContent, author: e.querySelector('.author').textContent })));
  return { method: 'scroll the page', quotes: quotes.length, authors: new Set(quotes.map((q) => q.author)).size, scrolls, requests, load_seconds: (loaded - t) / 1000, collect_seconds: (Date.now() - loaded) / 1000 };
}

// Way 2: catch the JSON request the page makes for its first screen, then call that endpoint
// yourself from inside the browser (same proxy, cookies and TLS fingerprint as the page),
// several pages at a time. No scrolling, no guessing when loading has finished.
async function byApi(page) {
  const t = Date.now();
  let requests = 0;
  page.on('request', () => { requests++; });
  const first = page.waitForResponse((r) => r.url().includes('/api/') && r.request().resourceType() === 'xhr', { timeout: 60000 });
  await page.goto(START, { timeout: 60000 });
  const response = await first;
  const loaded = Date.now();
  const endpoint = new URL(response.url());
  const quotes = [...(await response.json()).quotes]; // page 1 came for free
  const WAVE = 4; // one round trip through the proxy per wave, not per page
  let next = 2;
  for (let more = true; more; next += WAVE) {
    const urls = Array.from({ length: WAVE }, (_, i) => { endpoint.searchParams.set('page', String(next + i)); return endpoint.href; });
    const pages = await page.evaluate((us) => Promise.all(us.map((u) => fetch(u).then((r) => r.json()))), urls);
    for (const p of pages) quotes.push(...p.quotes);
    more = pages.every((p) => p.has_next);
  }
  return { method: 'call its JSON API', quotes: quotes.length, authors: new Set(quotes.map((q) => q.author.name)).size, api_pages: next - 1, endpoint: `${endpoint.pathname}?page=N`, requests, load_seconds: (loaded - t) / 1000, collect_seconds: (Date.now() - loaded) / 1000 };
}

const res = await fetch('https://starter.cdpfleet.com/chromium/session', {
  method: 'POST',
  headers: { 'x-api-key': KEY, 'content-type': 'application/json' },
  body: JSON.stringify({ proxy: process.env.PROXY_URL, headless: 'new' }),
});
if (!res.ok) throw new Error(`launch ${res.status} ${await res.text()}`);
const { wsUrl } = await res.json();
const browser = await chromium.connect(wsUrl, { headers: { 'x-api-key': KEY } });
try {
  const out = [];
  for (const way of [byScrolling, byApi]) {
    const page = await browser.newPage(); // a fresh tab per method, so request counts don't mix
    out.push(await way(page));
    await page.close();
  }
  console.log(JSON.stringify(out, null, 2));
} finally {
  await browser.close();
}

What we got

MethodQuotesAuthorsScrollsAPI pagesRequestsLoad (s)Collect (s)
scroll the page1005012—165.7618.785
call its JSON API10050—13194.6351.405

From the Node.js run on 2026-10-04. IP addresses are replaced with placeholders (203.0.113.x); equal addresses stay equal. The other languages produced the same findings.

Raw output (Node.js)
[
  {
    "method": "scroll the page",
    "quotes": 100,
    "authors": 50,
    "scrolls": 12,
    "requests": 16,
    "load_seconds": 5.761,
    "collect_seconds": 8.785
  },
  {
    "method": "call its JSON API",
    "quotes": 100,
    "authors": 50,
    "api_pages": 13,
    "endpoint": "/api/quotes?page=N",
    "requests": 19,
    "load_seconds": 4.635,
    "collect_seconds": 1.405
  }
]

Takeaways