Cases / #17 · 2026-10-04 · Easy
Infinite scroll, two ways: scroll like a user, or call the JSON the page calls
Collecting everything an endlessly scrolling page loads: by scrolling and watching the DOM, and by catching the page's own API request and paging through it from inside the browser.
Run on production on 2026-10-04: ✓ Node.js ✓ Python ✓ Java ✓ C# ✓ Go
The problem
Infinite-scroll pages have no "next" link and no end: content appears as you scroll, and the only way to know you've got everything is that scrolling stops producing more. The naive script scrolls, waits, counts, repeats — and has to guess how long to wait. But the page gets its data from somewhere: an XHR to a JSON endpoint, usually paginated, usually unauthenticated. Which approach is faster, which is more reliable, and how do you call that endpoint without losing the browser's proxy, cookies and TLS fingerprint?
What we used, and why
| What | Why |
|---|---|
chromium, headless: "new" | Nothing here depends on the browser; the cheapest session (1 thread) does. |
page.on("request") | Counts every request each method causes — page assets included — so the two can be compared. |
window.scrollTo(0, document.body.scrollHeight) in a loop | The user-like way: scroll, wait, count .quote elements, stop after three scrolls that added nothing. |
page.waitForResponse(xhr) | Catches the first JSON request the page makes for its own content, which gives us the endpoint, its parameters and page 1 of the data. |
page.evaluate(urls => Promise.all(urls.map(fetch))) | Calls the endpoint from inside the page: the browser's proxy, cookies, headers and TLS fingerprint, four pages per round trip. |
| A fresh tab per method | So request counts and timings don't mix. |
How it works
- Open the page and wait for the first quotes to render; record the load time separately from the collection time.
- Scroll: scroll to the bottom, wait 600 ms, count the quotes; stop when three scrolls in a row add nothing; read the quotes from the DOM.
- API: in a new tab, wait for the page's first XHR to
/api/…, take page 1 from its body, then fetch pages 2–5, 6–9, 10–13 in waves of four withfetchinside the page until a wave reports no next page. - Print, per method: quotes and distinct authors found, scrolls or API pages, requests made, load and collection seconds.
The code
The same program in five languages (also on GitHub, with the raw output). Set these environment variables first:
CDPFLEET_API_KEY— your API key (dashboard → API keys)PROXY_URL— your proxy, e.g.http://user:[email protected]:8000
// npm install [email protected]
// env: CDPFLEET_API_KEY, PROXY_URL
import { chromium } from 'playwright';
const KEY = process.env.CDPFLEET_API_KEY;
const START = 'https://quotes.toscrape.com/scroll'; // loads 10 quotes per screen, 100 in all
// Way 1: behave like a user — scroll to the bottom until nothing more appears, then read the DOM.
async function byScrolling(page) {
const t = Date.now();
let requests = 0;
page.on('request', () => { requests++; });
await page.goto(START, { timeout: 60000 });
await page.locator('.quote').first().waitFor({ timeout: 60000 });
const loaded = Date.now();
let count = 0;
let scrolls = 0;
let stale = 0;
while (stale < 3) {
await page.evaluate(() => window.scrollTo(0, document.body.scrollHeight));
scrolls++;
await page.waitForTimeout(600);
const now = await page.locator('.quote').count();
stale = now > count ? 0 : stale + 1;
count = now;
}
const quotes = await page.$$eval('.quote', (els) => els.map((e) => ({ text: e.querySelector('.text').textContent, author: e.querySelector('.author').textContent })));
return { method: 'scroll the page', quotes: quotes.length, authors: new Set(quotes.map((q) => q.author)).size, scrolls, requests, load_seconds: (loaded - t) / 1000, collect_seconds: (Date.now() - loaded) / 1000 };
}
// Way 2: catch the JSON request the page makes for its first screen, then call that endpoint
// yourself from inside the browser (same proxy, cookies and TLS fingerprint as the page),
// several pages at a time. No scrolling, no guessing when loading has finished.
async function byApi(page) {
const t = Date.now();
let requests = 0;
page.on('request', () => { requests++; });
const first = page.waitForResponse((r) => r.url().includes('/api/') && r.request().resourceType() === 'xhr', { timeout: 60000 });
await page.goto(START, { timeout: 60000 });
const response = await first;
const loaded = Date.now();
const endpoint = new URL(response.url());
const quotes = [...(await response.json()).quotes]; // page 1 came for free
const WAVE = 4; // one round trip through the proxy per wave, not per page
let next = 2;
for (let more = true; more; next += WAVE) {
const urls = Array.from({ length: WAVE }, (_, i) => { endpoint.searchParams.set('page', String(next + i)); return endpoint.href; });
const pages = await page.evaluate((us) => Promise.all(us.map((u) => fetch(u).then((r) => r.json()))), urls);
for (const p of pages) quotes.push(...p.quotes);
more = pages.every((p) => p.has_next);
}
return { method: 'call its JSON API', quotes: quotes.length, authors: new Set(quotes.map((q) => q.author.name)).size, api_pages: next - 1, endpoint: `${endpoint.pathname}?page=N`, requests, load_seconds: (loaded - t) / 1000, collect_seconds: (Date.now() - loaded) / 1000 };
}
const res = await fetch('https://starter.cdpfleet.com/chromium/session', {
method: 'POST',
headers: { 'x-api-key': KEY, 'content-type': 'application/json' },
body: JSON.stringify({ proxy: process.env.PROXY_URL, headless: 'new' }),
});
if (!res.ok) throw new Error(`launch ${res.status} ${await res.text()}`);
const { wsUrl } = await res.json();
const browser = await chromium.connect(wsUrl, { headers: { 'x-api-key': KEY } });
try {
const out = [];
for (const way of [byScrolling, byApi]) {
const page = await browser.newPage(); // a fresh tab per method, so request counts don't mix
out.push(await way(page));
await page.close();
}
console.log(JSON.stringify(out, null, 2));
} finally {
await browser.close();
}
# pip install playwright==1.60.0 requests
# env: CDPFLEET_API_KEY, PROXY_URL
import json
import os
import time
from urllib.parse import parse_qsl, urlencode, urlparse
import requests
from playwright.sync_api import sync_playwright
KEY = os.environ["CDPFLEET_API_KEY"]
START = "https://quotes.toscrape.com/scroll" # loads 10 quotes per screen, 100 in all
QUOTES_JS = "els => els.map(e => ({ text: e.querySelector('.text').textContent, author: e.querySelector('.author').textContent }))"
def by_scrolling(page):
"""Way 1: behave like a user — scroll to the bottom until nothing more appears, then read the DOM."""
t = time.time()
requests_ = {"n": 0}
page.on("request", lambda _: requests_.__setitem__("n", requests_["n"] + 1))
page.goto(START, timeout=60000)
page.locator(".quote").first.wait_for(timeout=60000)
loaded = time.time()
count = scrolls = stale = 0
while stale < 3:
page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
scrolls += 1
page.wait_for_timeout(600)
now = page.locator(".quote").count()
stale = 0 if now > count else stale + 1
count = now
quotes = page.eval_on_selector_all(".quote", QUOTES_JS)
done = time.time()
return {"method": "scroll the page", "quotes": len(quotes), "authors": len({q["author"] for q in quotes}),
"scrolls": scrolls, "requests": requests_["n"],
"load_seconds": round(loaded - t, 3), "collect_seconds": round(done - loaded, 3)}
def by_api(page):
"""Way 2: catch the JSON request the page makes for its first screen, then call that endpoint
yourself from inside the browser (same proxy, cookies and TLS fingerprint as the page),
several pages at a time. No scrolling, no guessing when loading has finished."""
t = time.time()
requests_ = {"n": 0}
page.on("request", lambda _: requests_.__setitem__("n", requests_["n"] + 1))
with page.expect_response(lambda r: "/api/" in r.url and r.request.resource_type == "xhr", timeout=60000) as first:
page.goto(START, timeout=60000)
response = first.value
loaded = time.time()
endpoint = urlparse(response.url)
params = dict(parse_qsl(endpoint.query))
quotes = list(response.json()["quotes"]) # page 1 came for free
wave = 4 # one round trip through the proxy per wave, not per page
nxt = 2
more = True
while more:
urls = [endpoint._replace(query=urlencode({**params, "page": str(n)})).geturl() for n in range(nxt, nxt + wave)]
pages = page.evaluate("us => Promise.all(us.map(u => fetch(u).then(r => r.json())))", urls)
for p in pages:
quotes.extend(p["quotes"])
more = all(p["has_next"] for p in pages)
nxt += wave
done = time.time()
return {"method": "call its JSON API", "quotes": len(quotes), "authors": len({q["author"]["name"] for q in quotes}),
"api_pages": nxt - 1, "endpoint": f"{endpoint.path}?page=N", "requests": requests_["n"],
"load_seconds": round(loaded - t, 3), "collect_seconds": round(done - loaded, 3)}
res = requests.post("https://starter.cdpfleet.com/chromium/session", headers={"x-api-key": KEY},
json={"proxy": os.environ["PROXY_URL"], "headless": "new"}, timeout=60)
if not res.ok:
raise SystemExit(f"launch {res.status_code} {res.text}")
with sync_playwright() as p:
browser = p.chromium.connect(res.json()["wsUrl"], headers={"x-api-key": KEY})
try:
out = []
for way in (by_scrolling, by_api):
page = browser.new_page() # a fresh tab per method, so request counts don't mix
out.append(way(page))
page.close()
print(json.dumps(out, indent=2, ensure_ascii=False))
finally:
browser.close()
// Maven: com.microsoft.playwright:playwright:1.60.0, com.google.code.gson:gson:2.11.0
// Run with PLAYWRIGHT_SKIP_BROWSER_DOWNLOAD=1. env: CDPFLEET_API_KEY, PROXY_URL
import com.google.gson.*;
import com.microsoft.playwright.*;
import java.net.URI;
import java.net.URLEncoder;
import java.net.http.*;
import java.nio.charset.StandardCharsets;
import java.util.*;
public class Main {
static final String KEY = System.getenv("CDPFLEET_API_KEY");
static final String START = "https://quotes.toscrape.com/scroll"; // loads 10 quotes per screen, 100 in all
static final String QUOTES_JS = "els => els.map(e => ({ text: e.querySelector('.text').textContent, author: e.querySelector('.author').textContent }))";
static final Gson GSON = new Gson();
// Way 1: behave like a user — scroll to the bottom until nothing more appears, then read the DOM.
static JsonObject byScrolling(Page page) {
long t = System.currentTimeMillis();
int[] requests = {0};
page.onRequest(req -> requests[0]++);
page.navigate(START, new Page.NavigateOptions().setTimeout(60000));
page.locator(".quote").first().waitFor(new Locator.WaitForOptions().setTimeout(60000));
long loaded = System.currentTimeMillis();
int count = 0, scrolls = 0, stale = 0;
while (stale < 3) {
page.evaluate("window.scrollTo(0, document.body.scrollHeight)");
scrolls++;
page.waitForTimeout(600);
int now = page.locator(".quote").count();
stale = now > count ? 0 : stale + 1;
count = now;
}
JsonArray quotes = GSON.toJsonTree(page.evalOnSelectorAll(".quote", QUOTES_JS)).getAsJsonArray();
long done = System.currentTimeMillis();
Set<String> authors = new HashSet<>();
for (JsonElement q : quotes) authors.add(q.getAsJsonObject().get("author").getAsString());
JsonObject row = new JsonObject();
row.addProperty("method", "scroll the page");
row.addProperty("quotes", quotes.size());
row.addProperty("authors", authors.size());
row.addProperty("scrolls", scrolls);
row.addProperty("requests", requests[0]);
row.addProperty("load_seconds", (loaded - t) / 1000.0);
row.addProperty("collect_seconds", (done - loaded) / 1000.0);
return row;
}
// The endpoint's URL with its page parameter replaced.
static String withPage(URI endpoint, int n) {
Map<String, String> params = new LinkedHashMap<>();
if (endpoint.getQuery() != null) for (String kv : endpoint.getQuery().split("&")) {
String[] p = kv.split("=", 2);
params.put(p[0], p.length > 1 ? p[1] : "");
}
params.put("page", String.valueOf(n));
StringJoiner q = new StringJoiner("&");
params.forEach((k, v) -> q.add(URLEncoder.encode(k, StandardCharsets.UTF_8) + "=" + URLEncoder.encode(v, StandardCharsets.UTF_8)));
return endpoint.getScheme() + "://" + endpoint.getAuthority() + endpoint.getPath() + "?" + q;
}
// Way 2: catch the JSON request the page makes for its first screen, then call that endpoint
// yourself from inside the browser (same proxy, cookies and TLS fingerprint as the page),
// several pages at a time. No scrolling, no guessing when loading has finished.
static JsonObject byApi(Page page) {
long t = System.currentTimeMillis();
int[] requests = {0};
page.onRequest(req -> requests[0]++);
Response response = page.waitForResponse(r -> r.url().contains("/api/") && r.request().resourceType().equals("xhr"),
new Page.WaitForResponseOptions().setTimeout(60000),
() -> page.navigate(START, new Page.NavigateOptions().setTimeout(60000)));
long loaded = System.currentTimeMillis();
URI endpoint = URI.create(response.url());
JsonArray quotes = new JsonArray();
quotes.addAll(JsonParser.parseString(response.text()).getAsJsonObject().getAsJsonArray("quotes")); // page 1 came for free
int wave = 4; // one round trip through the proxy per wave, not per page
int next = 2;
for (boolean more = true; more; next += wave) {
List<String> urls = new ArrayList<>();
for (int i = 0; i < wave; i++) urls.add(withPage(endpoint, next + i));
JsonArray pages = GSON.toJsonTree(page.evaluate("us => Promise.all(us.map(u => fetch(u).then(r => r.json())))", urls)).getAsJsonArray();
more = true;
for (JsonElement p : pages) {
quotes.addAll(p.getAsJsonObject().getAsJsonArray("quotes"));
more &= p.getAsJsonObject().get("has_next").getAsBoolean();
}
}
long done = System.currentTimeMillis();
Set<String> authors = new HashSet<>();
for (JsonElement q : quotes) authors.add(q.getAsJsonObject().getAsJsonObject("author").get("name").getAsString());
JsonObject row = new JsonObject();
row.addProperty("method", "call its JSON API");
row.addProperty("quotes", quotes.size());
row.addProperty("authors", authors.size());
row.addProperty("api_pages", next - 1);
row.addProperty("endpoint", endpoint.getPath() + "?page=N");
row.addProperty("requests", requests[0]);
row.addProperty("load_seconds", (loaded - t) / 1000.0);
row.addProperty("collect_seconds", (done - loaded) / 1000.0);
return row;
}
public static void main(String[] args) throws Exception {
String body = "{\"proxy\": " + GSON.toJson(System.getenv("PROXY_URL")) + ", \"headless\": \"new\"}";
HttpResponse<String> res = HttpClient.newHttpClient().send(HttpRequest.newBuilder(URI.create("https://starter.cdpfleet.com/chromium/session"))
.header("x-api-key", KEY).header("content-type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString(body)).build(), HttpResponse.BodyHandlers.ofString());
if (res.statusCode() != 200) throw new RuntimeException("launch: " + res.statusCode() + " " + res.body());
String wsUrl = JsonParser.parseString(res.body()).getAsJsonObject().get("wsUrl").getAsString();
try (Playwright playwright = Playwright.create()) {
Browser browser = playwright.chromium().connect(wsUrl, new BrowserType.ConnectOptions().setHeaders(Map.of("x-api-key", KEY)));
try {
JsonArray out = new JsonArray();
Page page = browser.newPage(); // a fresh tab per method, so request counts don't mix
out.add(byScrolling(page));
page.close();
page = browser.newPage();
out.add(byApi(page));
page.close();
System.out.println(new GsonBuilder().setPrettyPrinting().disableHtmlEscaping().create().toJson(out));
} finally {
browser.close();
}
}
}
}
// dotnet add package Microsoft.Playwright --version 1.60.0
// env: CDPFLEET_API_KEY, PROXY_URL
using System.Diagnostics;
using System.Net.Http.Json;
using System.Text.Encodings.Web;
using System.Text.Json;
using System.Text.Json.Nodes;
using System.Web;
using Microsoft.Playwright;
const string Start = "https://quotes.toscrape.com/scroll"; // loads 10 quotes per screen, 100 in all
const string QuotesJs = "els => els.map(e => ({ text: e.querySelector('.text').textContent, author: e.querySelector('.author').textContent }))";
var key = Environment.GetEnvironmentVariable("CDPFLEET_API_KEY")!;
using var http = new HttpClient();
http.DefaultRequestHeaders.Add("x-api-key", key);
var res = await http.PostAsJsonAsync("https://starter.cdpfleet.com/chromium/session",
new { proxy = Environment.GetEnvironmentVariable("PROXY_URL"), headless = "new" });
if (!res.IsSuccessStatusCode) throw new Exception($"launch: {(int)res.StatusCode} {await res.Content.ReadAsStringAsync()}");
var wsUrl = (await res.Content.ReadFromJsonAsync<JsonElement>()).GetProperty("wsUrl").GetString()!;
// Way 1: behave like a user — scroll to the bottom until nothing more appears, then read the DOM.
static async Task<JsonObject> ByScrolling(IPage page)
{
var sw = Stopwatch.StartNew();
int requests = 0;
page.Request += (_, _) => Interlocked.Increment(ref requests);
await page.GotoAsync(Start, new() { Timeout = 60000 });
await page.Locator(".quote").First.WaitForAsync(new() { Timeout = 60000 });
var loaded = sw.ElapsedMilliseconds;
int count = 0, scrolls = 0, stale = 0;
while (stale < 3)
{
await page.EvaluateAsync("window.scrollTo(0, document.body.scrollHeight)");
scrolls++;
await page.WaitForTimeoutAsync(600);
var now = await page.Locator(".quote").CountAsync();
stale = now > count ? 0 : stale + 1;
count = now;
}
var quotes = await page.EvalOnSelectorAllAsync<JsonElement>(".quote", QuotesJs);
var done = sw.ElapsedMilliseconds;
var authors = quotes.EnumerateArray().Select(q => q.GetProperty("author").GetString()).ToHashSet();
return new JsonObject
{
["method"] = "scroll the page", ["quotes"] = quotes.GetArrayLength(), ["authors"] = authors.Count,
["scrolls"] = scrolls, ["requests"] = requests,
["load_seconds"] = loaded / 1000.0, ["collect_seconds"] = (done - loaded) / 1000.0,
};
}
// The endpoint's URL with its page parameter replaced.
static string WithPage(Uri endpoint, int n)
{
var query = HttpUtility.ParseQueryString(endpoint.Query);
query["page"] = n.ToString();
return $"{endpoint.GetLeftPart(UriPartial.Path)}?{query}";
}
// Way 2: catch the JSON request the page makes for its first screen, then call that endpoint
// yourself from inside the browser (same proxy, cookies and TLS fingerprint as the page),
// several pages at a time. No scrolling, no guessing when loading has finished.
static async Task<JsonObject> ByApi(IPage page)
{
var sw = Stopwatch.StartNew();
int requests = 0;
page.Request += (_, _) => Interlocked.Increment(ref requests);
var response = await page.RunAndWaitForResponseAsync(
async () => await page.GotoAsync(Start, new() { Timeout = 60000 }),
r => r.Url.Contains("/api/") && r.Request.ResourceType == "xhr",
new() { Timeout = 60000 });
var loaded = sw.ElapsedMilliseconds;
var endpoint = new Uri(response.Url);
var quotes = (await response.JsonAsync())!.Value.GetProperty("quotes").EnumerateArray().ToList(); // page 1 came for free
const int wave = 4; // one round trip through the proxy per wave, not per page
int next = 2;
for (bool more = true; more; next += wave)
{
var urls = Enumerable.Range(next, wave).Select(n => WithPage(endpoint, n)).ToArray();
var pages = await page.EvaluateAsync<JsonElement>("us => Promise.all(us.map(u => fetch(u).then(r => r.json())))", urls);
more = true;
foreach (var p in pages.EnumerateArray())
{
quotes.AddRange(p.GetProperty("quotes").EnumerateArray());
more &= p.GetProperty("has_next").GetBoolean();
}
}
var done = sw.ElapsedMilliseconds;
var authors = quotes.Select(q => q.GetProperty("author").GetProperty("name").GetString()).ToHashSet();
return new JsonObject
{
["method"] = "call its JSON API", ["quotes"] = quotes.Count, ["authors"] = authors.Count,
["api_pages"] = next - 1, ["endpoint"] = $"{endpoint.AbsolutePath}?page=N", ["requests"] = requests,
["load_seconds"] = loaded / 1000.0, ["collect_seconds"] = (done - loaded) / 1000.0,
};
}
using var playwright = await Playwright.CreateAsync();
var browser = await playwright.Chromium.ConnectAsync(wsUrl, new() { Headers = new Dictionary<string, string> { ["x-api-key"] = key } });
try
{
var outRows = new JsonArray();
foreach (var way in new Func<IPage, Task<JsonObject>>[] { ByScrolling, ByApi })
{
var page = await browser.NewPageAsync(); // a fresh tab per method, so request counts don't mix
outRows.Add(await way(page));
await page.CloseAsync();
}
Console.WriteLine(outRows.ToJsonString(new JsonSerializerOptions { WriteIndented = true, Encoder = JavaScriptEncoder.UnsafeRelaxedJsonEscaping }));
}
finally
{
await browser.CloseAsync();
}
// go get github.com/playwright-community/[email protected]
// Driver: build playwright-core 1.60.0 from npm and set PLAYWRIGHT_DRIVER_PATH (see /docs/quickstart).
// env: CDPFLEET_API_KEY, PROXY_URL
package main
import (
"bytes"
"encoding/json"
"fmt"
"io"
"log"
"math"
"net/http"
"net/url"
"os"
"strconv"
"strings"
"sync"
"time"
"github.com/playwright-community/playwright-go"
)
var key = os.Getenv("CDPFLEET_API_KEY")
const start = "https://quotes.toscrape.com/scroll" // loads 10 quotes per screen, 100 in all
const quotesJS = "els => els.map(e => ({ text: e.querySelector('.text').textContent, author: e.querySelector('.author').textContent }))"
func launch(name string, options map[string]any) (map[string]any, error) {
body, _ := json.Marshal(options)
req, _ := http.NewRequest("POST", "https://starter.cdpfleet.com/"+name+"/session", bytes.NewReader(body))
req.Header.Set("x-api-key", key)
req.Header.Set("content-type", "application/json")
res, err := http.DefaultClient.Do(req)
if err != nil {
return nil, err
}
defer res.Body.Close()
if res.StatusCode != http.StatusOK {
msg, _ := io.ReadAll(res.Body)
return nil, fmt.Errorf("launch %s: %s %s", name, res.Status, msg)
}
var session map[string]any
return session, json.NewDecoder(res.Body).Decode(&session)
}
type scrollRow struct {
Method string `json:"method"`
Quotes int `json:"quotes"`
Authors int `json:"authors"`
Scrolls int `json:"scrolls"`
Requests int `json:"requests"`
LoadSeconds float64 `json:"load_seconds"`
CollectSeconds float64 `json:"collect_seconds"`
}
type apiRow struct {
Method string `json:"method"`
Quotes int `json:"quotes"`
Authors int `json:"authors"`
APIPages int `json:"api_pages"`
Endpoint string `json:"endpoint"`
Requests int `json:"requests"`
LoadSeconds float64 `json:"load_seconds"`
CollectSeconds float64 `json:"collect_seconds"`
}
func seconds(d time.Duration) float64 { return math.Round(d.Seconds()*1000) / 1000 }
// Counts every request the tab makes.
func countRequests(page playwright.Page) func() int {
var mu sync.Mutex
n := 0
page.OnRequest(func(playwright.Request) { mu.Lock(); n++; mu.Unlock() })
return func() int { mu.Lock(); defer mu.Unlock(); return n }
}
// Way 1: behave like a user — scroll to the bottom until nothing more appears, then read the DOM.
func byScrolling(page playwright.Page) scrollRow {
t := time.Now()
requests := countRequests(page)
if _, err := page.Goto(start, playwright.PageGotoOptions{Timeout: playwright.Float(60000)}); err != nil {
log.Fatal(err)
}
if err := page.Locator(".quote").First().WaitFor(playwright.LocatorWaitForOptions{Timeout: playwright.Float(60000)}); err != nil {
log.Fatal(err)
}
loaded := time.Now()
count, scrolls, stale := 0, 0, 0
for stale < 3 {
page.Evaluate("window.scrollTo(0, document.body.scrollHeight)")
scrolls++
page.WaitForTimeout(600)
now, _ := page.Locator(".quote").Count()
if now > count {
stale = 0
} else {
stale++
}
count = now
}
v, err := page.EvalOnSelectorAll(".quote", quotesJS)
if err != nil {
log.Fatal(err)
}
done := time.Now()
quotes := v.([]any)
authors := map[string]bool{}
for _, q := range quotes {
authors[q.(map[string]any)["author"].(string)] = true
}
return scrollRow{Method: "scroll the page", Quotes: len(quotes), Authors: len(authors), Scrolls: scrolls, Requests: requests(),
LoadSeconds: seconds(loaded.Sub(t)), CollectSeconds: seconds(done.Sub(loaded))}
}
// Way 2: catch the JSON request the page makes for its first screen, then call that endpoint
// yourself from inside the browser (same proxy, cookies and TLS fingerprint as the page),
// several pages at a time. No scrolling, no guessing when loading has finished.
func byAPI(page playwright.Page) apiRow {
t := time.Now()
requests := countRequests(page)
// playwright-go matches responses by URL only; the page's one /api/ request is its XHR.
response, err := page.ExpectResponse(func(u string) bool { return strings.Contains(u, "/api/") }, func() error {
_, err := page.Goto(start, playwright.PageGotoOptions{Timeout: playwright.Float(60000)})
return err
}, playwright.PageExpectResponseOptions{Timeout: playwright.Float(60000)})
if err != nil {
log.Fatal(err)
}
loaded := time.Now()
endpoint, _ := url.Parse(response.URL())
var first map[string]any
if err := response.JSON(&first); err != nil {
log.Fatal(err)
}
quotes := append([]any{}, first["quotes"].([]any)...) // page 1 came for free
const wave = 4 // one round trip through the proxy per wave, not per page
next := 2
for more := true; more; next += wave {
urls := make([]string, wave)
for i := range urls {
u := *endpoint
q := u.Query()
q.Set("page", strconv.Itoa(next+i))
u.RawQuery = q.Encode()
urls[i] = u.String()
}
v, err := page.Evaluate("us => Promise.all(us.map(u => fetch(u).then(r => r.json())))", urls)
if err != nil {
log.Fatal(err)
}
more = true
for _, p := range v.([]any) {
page := p.(map[string]any)
quotes = append(quotes, page["quotes"].([]any)...)
more = more && page["has_next"].(bool)
}
}
done := time.Now()
authors := map[string]bool{}
for _, q := range quotes {
authors[q.(map[string]any)["author"].(map[string]any)["name"].(string)] = true
}
return apiRow{Method: "call its JSON API", Quotes: len(quotes), Authors: len(authors), APIPages: next - 1,
Endpoint: endpoint.Path + "?page=N", Requests: requests(),
LoadSeconds: seconds(loaded.Sub(t)), CollectSeconds: seconds(done.Sub(loaded))}
}
func main() {
session, err := launch("chromium", map[string]any{"proxy": os.Getenv("PROXY_URL"), "headless": "new"})
if err != nil {
log.Fatal(err)
}
pw, err := playwright.Run(&playwright.RunOptions{SkipInstallBrowsers: true})
if err != nil {
log.Fatal(err)
}
defer pw.Stop()
browser, err := pw.Chromium.Connect(session["wsUrl"].(string), playwright.BrowserTypeConnectOptions{Headers: map[string]string{"x-api-key": key}})
if err != nil {
log.Fatal(err)
}
defer browser.Close()
var out []any
page, _ := browser.NewPage() // a fresh tab per method, so request counts don't mix
out = append(out, byScrolling(page))
page.Close()
page, _ = browser.NewPage()
out = append(out, byAPI(page))
page.Close()
var buf bytes.Buffer
enc := json.NewEncoder(&buf)
enc.SetEscapeHTML(false)
enc.SetIndent("", " ")
enc.Encode(out)
fmt.Print(buf.String())
}
What we got
| Method | Quotes | Authors | Scrolls | API pages | Requests | Load (s) | Collect (s) |
|---|---|---|---|---|---|---|---|
| scroll the page | 100 | 50 | 12 | — | 16 | 5.761 | 8.785 |
| call its JSON API | 100 | 50 | — | 13 | 19 | 4.635 | 1.405 |
From the Node.js run on 2026-10-04. IP addresses are replaced with placeholders (203.0.113.x); equal addresses stay equal. The other languages produced the same findings.
Raw output (Node.js)
[
{
"method": "scroll the page",
"quotes": 100,
"authors": 50,
"scrolls": 12,
"requests": 16,
"load_seconds": 5.761,
"collect_seconds": 8.785
},
{
"method": "call its JSON API",
"quotes": 100,
"authors": 50,
"api_pages": 13,
"endpoint": "/api/quotes?page=N",
"requests": 19,
"load_seconds": 4.635,
"collect_seconds": 1.405
}
]Takeaways
- Both methods find the same 100 quotes by 50 authors, so the JSON route loses nothing — and gains structure: the API returns authors as objects with slugs and tags as arrays, which the DOM route would have to parse.
- Collecting through the API was faster once the page had loaded — three round trips through the proxy for twelve pages instead of a run of scroll-and-wait cycles (5 to 15 across our runs: the faster the proxy, the more pages one cycle happens to catch). Through a residential proxy every round trip costs real time, so fetch pages in waves, not one by one: our first attempt, one request at a time, was slower than scrolling.
- Scrolling needs a stopping rule, and the rule costs time or correctness: "three scrolls that added nothing" means three waits after the real end on every run — and on a slow proxy moment it can still stop early, as one of our C# runs did at 10 quotes. The API tells you (
has_next: false). - Fetching from inside the page keeps you indistinguishable from the page: the proxy, the cookies, the
Accept/Originheaders and the TLS fingerprint are the browser's. A fetch from your script's own HTTP client would have none of them. - The last wave overshoots (pages 11–13 came back empty) — a cheap price for parallelism; size the wave to the latency you see.