Cases / #10 · 2026-09-30 · Hard
A thread-aware worker pool: one session per page vs. reusing sessions
24 pages on 6 parallel workers, two strategies, with the retries the API asks for. Reusing sessions is faster and costs a fraction of the thread time.
Run on production on 2026-09-30: ✓ Node.js ✓ Python ✓ Java ✓ C# ✓ Go
The problem
The simplest scraper launches a browser per URL. At scale that means paying the launch every time and holding a thread while the browser starts and stops. The alternative is a pool of long-lived sessions. How big is the difference — and how do you size the pool and handle the API's back-pressure properly?
What we used, and why
| What | Why |
|---|---|
GET https://cdpfleet.com/v1/me | Reads your plan, so the pool never has more workers than you have threads. |
| Launch retries | 429 (thread limit, launch rate) and 503 (fleet momentarily busy) carry Retry-After; wait that long and retry. Quota errors (daily_quota_exhausted, hourly_quota_exhausted) are not retried. |
| Page retries | Residential proxies occasionally drop a tunnel (ERR_TUNNEL_CONNECTION_FAILED): retry the navigation, up to three times. |
| Relaunch on a failed connect | Rarely, the server holding a new browser doesn't answer the WebSocket (502 upstream_unreachable). Don't reconnect to the same wsUrl — launch a fresh session. The failed one isn't billed. |
headless: true | One thread per worker — the cheapest mode for plain page loads. |
Wikipedia Special:Random | A cheap, always-different page for the workload. |
How it works
- Size the pool:
min(threads, 6)workers. - Strategy A: every page gets its own session (launch → connect → load → close).
- Strategy B: every worker opens one session and loads pages until the queue is empty.
- Compare wall time, launches and billed thread-seconds (the time sessions were open).
The code
The same program in five languages (also on GitHub, with the raw output). Set these environment variables first:
CDPFLEET_API_KEY— your API key (dashboard → API keys)PROXY_URL— your proxy, e.g.http://user:[email protected]:8000
// npm install [email protected]
// env: CDPFLEET_API_KEY, PROXY_URL
import { chromium } from 'playwright';
const KEY = process.env.CDPFLEET_API_KEY;
const PAGES = 24;
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
// Launch with the retries the API asks for: 429 (thread limit, launch rate) and 503
// (fleet momentarily busy) carry Retry-After.
async function launch(stats) {
for (let attempt = 1; ; attempt++) {
const t = Date.now();
const res = await fetch('https://starter.cdpfleet.com/chromium/session', {
method: 'POST',
headers: { 'x-api-key': KEY, 'content-type': 'application/json' },
body: JSON.stringify({ proxy: process.env.PROXY_URL, headless: true }),
});
if (res.ok) return { ...(await res.json()), launchMs: Date.now() - t };
const body = await res.json().catch(() => ({}));
const retryable = res.status === 503 || (res.status === 429 && !/quota/.test(body.error));
if (!retryable || attempt === 10) throw new Error(`launch: ${res.status} ${body.error}`);
stats.retries[body.error] = (stats.retries[body.error] || 0) + 1;
await sleep(Number(res.headers.get('retry-after') || 2) * 1000);
}
}
// Residential proxies drop a tunnel now and then (ERR_TUNNEL_CONNECTION_FAILED): retry.
async function scrape(page, stats) {
for (let attempt = 1; ; attempt++) {
try {
await page.goto('https://en.wikipedia.org/wiki/Special:Random', { timeout: 30000 });
return page.title();
} catch (err) {
stats.pageRetries++;
if (attempt === 3) return `(failed: ${err.message.split('\n')[0]})`;
}
}
}
// Runs `jobs` pages on `workers` parallel workers; each worker either opens one session
// and reuses it, or opens a new session for every page.
async function run(workers, reuse) {
const stats = { launches: 0, launchMs: 0, connectMs: 0, sessionSeconds: 0, titles: [], retries: {}, pageRetries: 0, connectFailures: 0 };
let next = 0;
// If the connect fails (rare: the server holding the browser didn't answer), don't
// reconnect to the same wsUrl — launch a fresh session. Such sessions aren't billed.
const openSession = async () => {
for (let attempt = 1; ; attempt++) {
const s = await launch(stats);
stats.launches++;
stats.launchMs += s.launchMs;
const t = Date.now();
try {
const browser = await chromium.connect(s.wsUrl, { headers: { 'x-api-key': KEY } });
stats.connectMs += Date.now() - t;
return { browser, started: t };
} catch (err) {
stats.connectFailures++;
if (attempt === 3) throw err;
}
}
};
const closeSession = async ({ browser, started }) => {
await browser.close();
stats.sessionSeconds += (Date.now() - started) / 1000;
};
const worker = async () => {
if (reuse) {
const session = await openSession();
try {
const page = await session.browser.newPage();
while (next < PAGES) { next++; stats.titles.push(await scrape(page, stats)); }
} finally {
await closeSession(session);
}
return;
}
while (next < PAGES) {
next++;
const session = await openSession();
try {
stats.titles.push(await scrape(await session.browser.newPage(), stats));
} finally {
await closeSession(session);
}
}
};
const t = Date.now();
await Promise.all(Array.from({ length: workers }, worker));
return {
strategy: reuse ? 'reuse one session per worker' : 'new session per page',
pages: stats.titles.length,
wall_seconds: Math.round((Date.now() - t) / 100) / 10,
launches: stats.launches,
avg_launch_ms: Math.round(stats.launchMs / stats.launches),
avg_connect_ms: Math.round(stats.connectMs / stats.launches),
launch_retries: stats.retries,
connect_failures: stats.connectFailures,
page_retries: stats.pageRetries,
billed_thread_seconds: Math.round(stats.sessionSeconds),
sample_titles: stats.titles.slice(0, 3),
};
}
// Size the pool from the plan: never more workers than threads.
const me = await (await fetch('https://cdpfleet.com/v1/me', { headers: { 'x-api-key': KEY } })).json();
const workers = Math.min(me.subscription.threads, 6);
console.log(JSON.stringify({
plan_threads: me.subscription.threads,
workers,
results: [await run(workers, false), await run(workers, true)],
}, null, 2));
# pip install playwright==1.60.0 aiohttp
# env: CDPFLEET_API_KEY, PROXY_URL
import asyncio
import json
import os
import time
import aiohttp
from playwright.async_api import async_playwright
KEY = os.environ["CDPFLEET_API_KEY"]
PAGES = 24
async def launch(http, stats):
"""Launch with the retries the API asks for: 429 (thread limit, launch rate) and 503
(fleet momentarily busy) carry Retry-After."""
for attempt in range(1, 11):
t = time.time()
async with http.post("https://starter.cdpfleet.com/chromium/session", headers={"x-api-key": KEY},
json={"proxy": os.environ["PROXY_URL"], "headless": True}) as res:
body = await res.json(content_type=None)
if res.status == 200:
return body, (time.time() - t) * 1000
retryable = res.status == 503 or (res.status == 429 and "quota" not in body.get("error", ""))
if not retryable or attempt == 10:
raise RuntimeError(f"launch: {res.status} {body.get('error')}")
stats["retries"][body.get("error")] = stats["retries"].get(body.get("error"), 0) + 1
await asyncio.sleep(float(res.headers.get("retry-after", 2)))
async def scrape(page, stats):
# Residential proxies drop a tunnel now and then (ERR_TUNNEL_CONNECTION_FAILED): retry.
for attempt in (1, 2, 3):
try:
await page.goto("https://en.wikipedia.org/wiki/Special:Random", timeout=30000)
return await page.title()
except Exception as err:
stats["page_retries"] += 1
if attempt == 3:
return f"(failed: {str(err).splitlines()[0]})"
async def run(p, http, workers, reuse):
"""Runs PAGES pages on `workers` parallel workers; each worker either opens one session
and reuses it, or opens a new session for every page."""
stats = {"launches": 0, "launch_ms": 0, "connect_ms": 0, "session_s": 0, "titles": [], "retries": {}, "page_retries": 0, "connect_failures": 0}
queue = asyncio.Queue()
for i in range(PAGES):
queue.put_nowait(i)
async def open_session():
# If the connect fails (rare: the server holding the browser didn't answer), don't
# reconnect to the same wsUrl — launch a fresh session. Such sessions aren't billed.
for attempt in (1, 2, 3):
session, launch_ms = await launch(http, stats)
stats["launches"] += 1
stats["launch_ms"] += launch_ms
t = time.time()
try:
browser = await p.chromium.connect(session["wsUrl"], headers={"x-api-key": KEY})
except Exception:
stats["connect_failures"] += 1
if attempt == 3:
raise
continue
stats["connect_ms"] += (time.time() - t) * 1000
return browser, t
async def close_session(browser, started):
await browser.close()
stats["session_s"] += time.time() - started
async def worker():
if reuse:
browser, started = await open_session()
try:
page = await browser.new_page()
while not queue.empty():
queue.get_nowait()
stats["titles"].append(await scrape(page, stats))
finally:
await close_session(browser, started)
return
while not queue.empty():
queue.get_nowait()
browser, started = await open_session()
try:
stats["titles"].append(await scrape(await browser.new_page(), stats))
finally:
await close_session(browser, started)
t = time.time()
await asyncio.gather(*(worker() for _ in range(workers)))
return {
"strategy": "reuse one session per worker" if reuse else "new session per page",
"pages": len(stats["titles"]),
"wall_seconds": round(time.time() - t, 1),
"launches": stats["launches"],
"avg_launch_ms": round(stats["launch_ms"] / stats["launches"]),
"avg_connect_ms": round(stats["connect_ms"] / stats["launches"]),
"launch_retries": stats["retries"],
"connect_failures": stats["connect_failures"],
"page_retries": stats["page_retries"],
"billed_thread_seconds": round(stats["session_s"]),
"sample_titles": stats["titles"][:3],
}
async def main():
async with aiohttp.ClientSession() as http, async_playwright() as p:
# Size the pool from the plan: never more workers than threads.
async with http.get("https://cdpfleet.com/v1/me", headers={"x-api-key": KEY}) as r:
threads = (await r.json())["subscription"]["threads"]
workers = min(threads, 6)
results = [await run(p, http, workers, False), await run(p, http, workers, True)]
print(json.dumps({"plan_threads": threads, "workers": workers, "results": results}, indent=2))
asyncio.run(main())
// Maven: com.microsoft.playwright:playwright:1.60.0, com.google.code.gson:gson:2.11.0
// Run with PLAYWRIGHT_SKIP_BROWSER_DOWNLOAD=1. env: CDPFLEET_API_KEY, PROXY_URL
import com.google.gson.*;
import com.microsoft.playwright.*;
import java.net.URI;
import java.net.http.*;
import java.util.*;
import java.util.concurrent.*;
import java.util.concurrent.atomic.*;
public class Main {
static final String KEY = System.getenv("CDPFLEET_API_KEY");
static final HttpClient HTTP = HttpClient.newHttpClient();
static final int PAGES = 24;
static class Stats {
final AtomicInteger launches = new AtomicInteger(), pageRetries = new AtomicInteger(), connectFailures = new AtomicInteger();
final AtomicLong launchMs = new AtomicLong(), connectMs = new AtomicLong(), sessionMs = new AtomicLong();
final Map<String, Integer> retries = new ConcurrentHashMap<>();
final List<String> titles = Collections.synchronizedList(new ArrayList<>());
}
// Launch with the retries the API asks for: 429 (thread limit, launch rate) and 503
// (fleet momentarily busy) carry Retry-After.
static String launch(Stats stats) throws Exception {
String body = "{\"proxy\": " + new Gson().toJson(System.getenv("PROXY_URL")) + ", \"headless\": true}";
for (int attempt = 1; ; attempt++) {
long t = System.currentTimeMillis();
HttpResponse<String> res = HTTP.send(HttpRequest.newBuilder(URI.create("https://starter.cdpfleet.com/chromium/session"))
.header("x-api-key", KEY).header("content-type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString(body)).build(), HttpResponse.BodyHandlers.ofString());
JsonObject json = JsonParser.parseString(res.body()).getAsJsonObject();
if (res.statusCode() == 200) {
stats.launches.incrementAndGet();
stats.launchMs.addAndGet(System.currentTimeMillis() - t);
return json.get("wsUrl").getAsString();
}
String error = json.has("error") ? json.get("error").getAsString() : "";
boolean retryable = res.statusCode() == 503 || (res.statusCode() == 429 && !error.contains("quota"));
if (!retryable || attempt == 10) throw new RuntimeException("launch: " + res.statusCode() + " " + error);
stats.retries.merge(error, 1, Integer::sum);
Thread.sleep(Long.parseLong(res.headers().firstValue("retry-after").orElse("2")) * 1000);
}
}
// Residential proxies drop a tunnel now and then (ERR_TUNNEL_CONNECTION_FAILED): retry.
static String scrape(Page page, Stats stats) {
for (int attempt = 1; ; attempt++) {
try {
page.navigate("https://en.wikipedia.org/wiki/Special:Random", new Page.NavigateOptions().setTimeout(30000));
return page.title();
} catch (PlaywrightException err) {
stats.pageRetries.incrementAndGet();
if (attempt == 3) return "(failed: " + err.getMessage().split("\n")[0] + ")";
}
}
}
// One worker thread. Playwright objects are per thread, so each worker has its own.
static void worker(boolean reuse, AtomicInteger next, Stats stats) throws Exception {
try (Playwright playwright = Playwright.create()) {
Browser browser = null;
long started = 0;
Page page = null;
while (next.getAndIncrement() < PAGES) {
if (browser == null) {
// If the connect fails (rare: the server holding the browser didn't answer), don't
// reconnect to the same wsUrl — launch a fresh session. Such sessions aren't billed.
for (int attempt = 1; browser == null; attempt++) {
String wsUrl = launch(stats);
started = System.currentTimeMillis();
try {
browser = playwright.chromium().connect(wsUrl, new BrowserType.ConnectOptions().setHeaders(Map.of("x-api-key", KEY)));
} catch (PlaywrightException err) {
stats.connectFailures.incrementAndGet();
if (attempt == 3) throw err;
}
}
stats.connectMs.addAndGet(System.currentTimeMillis() - started);
page = browser.newPage();
}
try {
stats.titles.add(scrape(page, stats));
} finally {
if (!reuse) {
browser.close();
stats.sessionMs.addAndGet(System.currentTimeMillis() - started);
browser = null;
}
}
}
if (browser != null) {
browser.close();
stats.sessionMs.addAndGet(System.currentTimeMillis() - started);
}
}
}
// Runs PAGES pages on `workers` parallel workers; each worker either opens one session
// and reuses it, or opens a new session for every page.
static JsonObject run(int workers, boolean reuse) throws Exception {
Stats stats = new Stats();
AtomicInteger next = new AtomicInteger();
long t = System.currentTimeMillis();
ExecutorService pool = Executors.newFixedThreadPool(workers);
List<Future<?>> futures = new ArrayList<>();
for (int i = 0; i < workers; i++) futures.add(pool.submit(() -> { worker(reuse, next, stats); return null; }));
for (Future<?> f : futures) f.get();
pool.shutdown();
Gson gson = new Gson();
JsonObject out = new JsonObject();
out.addProperty("strategy", reuse ? "reuse one session per worker" : "new session per page");
out.addProperty("pages", stats.titles.size());
out.addProperty("wall_seconds", Math.round((System.currentTimeMillis() - t) / 100.0) / 10.0);
out.addProperty("launches", stats.launches.get());
out.addProperty("avg_launch_ms", stats.launchMs.get() / stats.launches.get());
out.addProperty("avg_connect_ms", stats.connectMs.get() / stats.launches.get());
out.add("launch_retries", gson.toJsonTree(stats.retries));
out.addProperty("connect_failures", stats.connectFailures.get());
out.addProperty("page_retries", stats.pageRetries.get());
out.addProperty("billed_thread_seconds", Math.round(stats.sessionMs.get() / 1000.0));
out.add("sample_titles", gson.toJsonTree(stats.titles.subList(0, Math.min(3, stats.titles.size()))));
return out;
}
public static void main(String[] args) throws Exception {
// Size the pool from the plan: never more workers than threads.
HttpResponse<String> me = HTTP.send(HttpRequest.newBuilder(URI.create("https://cdpfleet.com/v1/me")).header("x-api-key", KEY).build(),
HttpResponse.BodyHandlers.ofString());
int threads = JsonParser.parseString(me.body()).getAsJsonObject().getAsJsonObject("subscription").get("threads").getAsInt();
int workers = Math.min(threads, 6);
JsonObject out = new JsonObject();
out.addProperty("plan_threads", threads);
out.addProperty("workers", workers);
JsonArray results = new JsonArray();
results.add(run(workers, false));
results.add(run(workers, true));
out.add("results", results);
System.out.println(new GsonBuilder().setPrettyPrinting().disableHtmlEscaping().create().toJson(out));
}
}
// dotnet add package Microsoft.Playwright --version 1.60.0
// env: CDPFLEET_API_KEY, PROXY_URL
using System.Collections.Concurrent;
using System.Diagnostics;
using System.Net.Http.Json;
using System.Text.Encodings.Web;
using System.Text.Json;
using System.Text.Json.Nodes;
using Microsoft.Playwright;
const int Pages = 24;
var key = Environment.GetEnvironmentVariable("CDPFLEET_API_KEY")!;
using var http = new HttpClient();
http.DefaultRequestHeaders.Add("x-api-key", key);
using var playwright = await Playwright.CreateAsync();
// Launch with the retries the API asks for: 429 (thread limit, launch rate) and 503
// (fleet momentarily busy) carry Retry-After.
async Task<(string WsUrl, long Ms)> Launch(ConcurrentDictionary<string, int> retries)
{
for (var attempt = 1; ; attempt++)
{
var sw = Stopwatch.StartNew();
var res = await http.PostAsJsonAsync("https://starter.cdpfleet.com/chromium/session",
new { proxy = Environment.GetEnvironmentVariable("PROXY_URL"), headless = true });
var body = await res.Content.ReadFromJsonAsync<JsonElement>();
if (res.IsSuccessStatusCode) return (body.GetProperty("wsUrl").GetString()!, sw.ElapsedMilliseconds);
var error = body.TryGetProperty("error", out var e) ? e.GetString() ?? "" : "";
var retryable = (int)res.StatusCode == 503 || ((int)res.StatusCode == 429 && !error.Contains("quota"));
if (!retryable || attempt == 10) throw new Exception($"launch: {(int)res.StatusCode} {error}");
retries.AddOrUpdate(error, 1, (_, n) => n + 1);
await Task.Delay(TimeSpan.FromSeconds(res.Headers.RetryAfter?.Delta?.TotalSeconds ?? 2));
}
}
// Residential proxies drop a tunnel now and then (ERR_TUNNEL_CONNECTION_FAILED): retry.
async Task<string> Scrape(IPage page, Action retried)
{
for (var attempt = 1; ; attempt++)
{
try
{
await page.GotoAsync("https://en.wikipedia.org/wiki/Special:Random", new() { Timeout = 30000 });
return await page.TitleAsync();
}
catch (Exception err) when (err is PlaywrightException or TimeoutException) // .NET times out with TimeoutException
{
retried();
if (attempt == 3) return $"(failed: {err.Message.Split('\n')[0]})";
}
}
}
// Runs Pages pages on `workers` parallel workers; each worker either opens one session
// and reuses it, or opens a new session for every page.
async Task<JsonObject> Run(int workers, bool reuse)
{
var titles = new ConcurrentQueue<string>();
var retries = new ConcurrentDictionary<string, int>();
long launches = 0, launchMs = 0, connectMs = 0, sessionMs = 0, pageRetries = 0, connectFailures = 0;
var next = -1;
// If the connect fails (rare: the server holding the browser didn't answer), don't
// reconnect to the same wsUrl — launch a fresh session. Such sessions aren't billed.
async Task<(IBrowser Browser, Stopwatch Clock)> Open()
{
for (var attempt = 1; ; attempt++)
{
var (ws, ms) = await Launch(retries);
Interlocked.Increment(ref launches);
Interlocked.Add(ref launchMs, ms);
var clock = Stopwatch.StartNew();
try
{
var browser = await playwright.Chromium.ConnectAsync(ws, new() { Headers = new Dictionary<string, string> { ["x-api-key"] = key } });
Interlocked.Add(ref connectMs, clock.ElapsedMilliseconds);
return (browser, clock);
}
catch (Exception err) when ((err is PlaywrightException or TimeoutException) && attempt < 3)
{
Interlocked.Increment(ref connectFailures);
}
}
}
async Task Close((IBrowser Browser, Stopwatch Clock) s)
{
await s.Browser.CloseAsync();
Interlocked.Add(ref sessionMs, s.Clock.ElapsedMilliseconds);
}
async Task Worker()
{
if (reuse)
{
var s = await Open();
try
{
var page = await s.Browser.NewPageAsync();
while (Interlocked.Increment(ref next) < Pages) titles.Enqueue(await Scrape(page, () => Interlocked.Increment(ref pageRetries)));
}
finally { await Close(s); }
return;
}
while (Interlocked.Increment(ref next) < Pages)
{
var s = await Open();
try { titles.Enqueue(await Scrape(await s.Browser.NewPageAsync(), () => Interlocked.Increment(ref pageRetries))); }
finally { await Close(s); }
}
}
var wall = Stopwatch.StartNew();
await Task.WhenAll(Enumerable.Range(0, workers).Select(_ => Worker()));
return new JsonObject
{
["strategy"] = reuse ? "reuse one session per worker" : "new session per page",
["pages"] = titles.Count,
["wall_seconds"] = Math.Round(wall.Elapsed.TotalSeconds, 1),
["launches"] = launches,
["avg_launch_ms"] = launchMs / launches,
["avg_connect_ms"] = connectMs / launches,
["launch_retries"] = new JsonObject(retries.Select(kv => KeyValuePair.Create(kv.Key, (JsonNode?)kv.Value))),
["connect_failures"] = connectFailures,
["page_retries"] = pageRetries,
["billed_thread_seconds"] = (long)Math.Round(sessionMs / 1000.0),
["sample_titles"] = new JsonArray(titles.Take(3).Select(t => (JsonNode?)t).ToArray()),
};
}
// Size the pool from the plan: never more workers than threads.
var me = JsonNode.Parse(await http.GetStringAsync("https://cdpfleet.com/v1/me"))!;
var threads = (int)me["subscription"]!["threads"]!;
var workerCount = Math.Min(threads, 6);
var result = new JsonObject
{
["plan_threads"] = threads,
["workers"] = workerCount,
["results"] = new JsonArray(await Run(workerCount, false), await Run(workerCount, true)),
};
Console.WriteLine(result.ToJsonString(new JsonSerializerOptions { WriteIndented = true, Encoder = JavaScriptEncoder.UnsafeRelaxedJsonEscaping }));
// go get github.com/playwright-community/[email protected]
// Driver: build playwright-core 1.60.0 from npm and set PLAYWRIGHT_DRIVER_PATH (see /docs/quickstart).
// env: CDPFLEET_API_KEY, PROXY_URL
package main
import (
"bytes"
"encoding/json"
"fmt"
"log"
"math"
"net/http"
"os"
"strconv"
"strings"
"sync"
"sync/atomic"
"time"
"github.com/playwright-community/playwright-go"
)
var key = os.Getenv("CDPFLEET_API_KEY")
const pages = 24
type stats struct {
mu sync.Mutex
launches, launchMs, connectMs, sessionMs, pageRetries, connectFailures int64
retries map[string]int
titles []string
}
// Launch with the retries the API asks for: 429 (thread limit, launch rate) and 503
// (fleet momentarily busy) carry Retry-After.
func launch(s *stats) (string, error) {
body, _ := json.Marshal(map[string]any{"proxy": os.Getenv("PROXY_URL"), "headless": true})
for attempt := 1; ; attempt++ {
t := time.Now()
req, _ := http.NewRequest("POST", "https://starter.cdpfleet.com/chromium/session", bytes.NewReader(body))
req.Header.Set("x-api-key", key)
req.Header.Set("content-type", "application/json")
res, err := http.DefaultClient.Do(req)
if err != nil {
return "", err
}
var out map[string]any
json.NewDecoder(res.Body).Decode(&out)
res.Body.Close()
if res.StatusCode == http.StatusOK {
atomic.AddInt64(&s.launches, 1)
atomic.AddInt64(&s.launchMs, time.Since(t).Milliseconds())
return out["wsUrl"].(string), nil
}
errName, _ := out["error"].(string)
retryable := res.StatusCode == 503 || (res.StatusCode == 429 && !strings.Contains(errName, "quota"))
if !retryable || attempt == 10 {
return "", fmt.Errorf("launch: %d %s", res.StatusCode, errName)
}
s.mu.Lock()
s.retries[errName]++
s.mu.Unlock()
wait, err := strconv.Atoi(res.Header.Get("retry-after"))
if err != nil {
wait = 2
}
time.Sleep(time.Duration(wait) * time.Second)
}
}
// Residential proxies drop a tunnel now and then (ERR_TUNNEL_CONNECTION_FAILED): retry.
func scrape(page playwright.Page, s *stats) string {
for attempt := 1; ; attempt++ {
_, err := page.Goto("https://en.wikipedia.org/wiki/Special:Random", playwright.PageGotoOptions{Timeout: playwright.Float(30000)})
if err == nil {
title, _ := page.Title()
return title
}
atomic.AddInt64(&s.pageRetries, 1)
if attempt == 3 {
return "(failed: " + strings.SplitN(err.Error(), "\n", 2)[0] + ")"
}
}
}
// Runs pages pages on workers parallel workers; each worker either opens one session
// and reuses it, or opens a new session for every page.
func run(pw *playwright.Playwright, workers int, reuse bool) map[string]any {
s := &stats{retries: map[string]int{}}
var next int64 = -1
// If the connect fails (rare: the server holding the browser didn't answer), don't
// reconnect to the same wsUrl: launch a fresh session. Such sessions aren't billed.
open := func() (playwright.Browser, time.Time) {
for attempt := 1; ; attempt++ {
ws, err := launch(s)
if err != nil {
log.Fatal(err)
}
t := time.Now()
browser, err := pw.Chromium.Connect(ws, playwright.BrowserTypeConnectOptions{Headers: map[string]string{"x-api-key": key}})
if err == nil {
atomic.AddInt64(&s.connectMs, time.Since(t).Milliseconds())
return browser, t
}
atomic.AddInt64(&s.connectFailures, 1)
if attempt == 3 {
log.Fatal(err)
}
}
}
closeSession := func(b playwright.Browser, started time.Time) {
b.Close()
atomic.AddInt64(&s.sessionMs, time.Since(started).Milliseconds())
}
add := func(title string) { s.mu.Lock(); s.titles = append(s.titles, title); s.mu.Unlock() }
worker := func() {
if reuse {
browser, started := open()
defer closeSession(browser, started)
page, _ := browser.NewPage()
for atomic.AddInt64(&next, 1) < pages {
add(scrape(page, s))
}
return
}
for atomic.AddInt64(&next, 1) < pages {
browser, started := open()
page, _ := browser.NewPage()
add(scrape(page, s))
closeSession(browser, started)
}
}
t := time.Now()
var wg sync.WaitGroup
for i := 0; i < workers; i++ {
wg.Add(1)
go func() { defer wg.Done(); worker() }()
}
wg.Wait()
strategy := "new session per page"
if reuse {
strategy = "reuse one session per worker"
}
return map[string]any{
"strategy": strategy,
"pages": len(s.titles),
"wall_seconds": math.Round(time.Since(t).Seconds()*10) / 10,
"launches": s.launches,
"avg_launch_ms": s.launchMs / s.launches,
"avg_connect_ms": s.connectMs / s.launches,
"launch_retries": s.retries,
"connect_failures": s.connectFailures,
"page_retries": s.pageRetries,
"billed_thread_seconds": int64(math.Round(float64(s.sessionMs) / 1000)),
"sample_titles": s.titles[:min(3, len(s.titles))],
}
}
func main() {
// Size the pool from the plan: never more workers than threads.
req, _ := http.NewRequest("GET", "https://cdpfleet.com/v1/me", nil)
req.Header.Set("x-api-key", key)
res, err := http.DefaultClient.Do(req)
if err != nil {
log.Fatal(err)
}
var me struct {
Subscription struct{ Threads int } `json:"subscription"`
}
json.NewDecoder(res.Body).Decode(&me)
res.Body.Close()
workers := min(me.Subscription.Threads, 6)
pw, err := playwright.Run(&playwright.RunOptions{SkipInstallBrowsers: true})
if err != nil {
log.Fatal(err)
}
defer pw.Stop()
out, _ := json.MarshalIndent(map[string]any{
"plan_threads": me.Subscription.Threads,
"workers": workers,
"results": []map[string]any{run(pw, workers, false), run(pw, workers, true)},
}, "", " ")
fmt.Println(string(out))
}
What we got
| Strategy | Pages | Wall time (s) | Launches | Launch (ms) | Billed thread-s | Connect failures | Page retries |
|---|---|---|---|---|---|---|---|
| new session per page | 24 | 49.6 | 24 | 310 | 177 | 0 | 0 |
| reuse one session per worker | 24 | 23 | 6 | 780 | 106 | 0 | 0 |
From the Node.js run on 2026-09-30. IP addresses are replaced with placeholders (203.0.113.x); equal addresses stay equal. The other languages produced the same findings.
Raw output (Node.js)
{
"plan_threads": 100,
"workers": 6,
"results": [
{
"strategy": "new session per page",
"pages": 24,
"wall_seconds": 49.6,
"launches": 24,
"avg_launch_ms": 310,
"avg_connect_ms": 84,
"launch_retries": {},
"connect_failures": 0,
"page_retries": 0,
"billed_thread_seconds": 177,
"sample_titles": [
"Jennie Litvack - Wikipedia",
"Athletics at the 1983 Pan American Games – Men's 1500 metres - Wikipedia",
"Cornelius Cruys - Wikipedia"
]
},
{
"strategy": "reuse one session per worker",
"pages": 24,
"wall_seconds": 23,
"launches": 6,
"avg_launch_ms": 780,
"avg_connect_ms": 71,
"launch_retries": {},
"connect_failures": 0,
"page_retries": 0,
"billed_thread_seconds": 106,
"sample_titles": [
"Pallichattambi - Wikipedia",
"Poliez-le-Grand - Wikipedia",
"Emiliano Mayola - Wikipedia"
]
}
]
}Takeaways
- Reusing sessions billed a third to half of the thread time for the same 24 pages, and in clean runs finished in about half the wall time. A single slow proxy tunnel can dominate either strategy — that's what
page_retriesis for. - Launching is quick (~0.3–0.5 s) — the cost is everything around it: connecting, the first navigation on a cold browser, and tearing it down, all of which you pay per page in strategy A.
- Size from the plan, not a constant: a pool larger than your threads just collects
429 threads_exceeded. - Retries are part of the design, not an afterthought: honour
Retry-Afterfor launches, relaunch on a failed connect, retry navigations on proxy errors — and count all three so a slow run explains itself. - Start a fresh session when you need a fresh identity (new cookies, fingerprint or proxy) — otherwise reuse.