Cases / #19 · 2026-10-05 · Easy
Fetch first, browser second: paying for JavaScript only when the page needs it
A plain HTTP GET through your proxy tells you in a second whether the data is in the HTML. Only the pages that render with JavaScript get a browser session — three pages, one decision rule, measured.
Run on production on 2026-10-05: ✓ Node.js ✓ Python ✓ Java ✓ C# ✓ Go
The problem
A browser is the expensive way to download HTML. Many pages — catalogues, listings, documentation — ship their data in the HTML and need no JavaScript at all; others are empty shells filled in by scripts, and some need scrolling or clicks on top. If you send every URL to a browser you pay thread-seconds for pages curl could have fetched; if you send none you silently get empty results from the JavaScript ones. The question is how to decide per page, cheaply, through the same proxy, without maintaining two code paths.
What we used, and why
| What | Why |
|---|---|
playwright.request.newContext({ proxy }) | Playwright's HTTP client, run locally with your proxy — no session, no thread, same exit IP as the browser would have. Available in all five languages. |
A browser-like User-Agent on the plain request | So the server answers the plain GET the way it answers the browser; otherwise some sites serve a different page to a bare client. |
| Counting the item markup in the raw HTML | The decision rule: if the HTML already contains the expected items (class="product_pod", class="quote"), there is nothing a browser would add. |
chromium, headless: "new", launched lazily | One session for all the pages that failed the rule; none at all if every page passed. |
locator(…).first().waitFor() then count() | In the browser, wait for the rendered items rather than for "load": JavaScript-rendered pages fire load before their content exists. |
How it works
- GET each page through the proxy with the plain client; record status, size, how many items the HTML contains and the time.
- Pages whose HTML has fewer items than expected are marked as needing a browser.
- Launch one headless Chromium session only if that list is non-empty; open each such page, wait for the items to render, count them, record the time.
- Print one row per page: what the HTML had, what the browser had, which path it needed.
The code
The same program in five languages (also on GitHub, with the raw output). Set these environment variables first:
CDPFLEET_API_KEY— your API key (dashboard → API keys)PROXY_URL— your proxy, e.g.http://user:[email protected]:8000
// npm install [email protected]
// env: CDPFLEET_API_KEY, PROXY_URL
import { chromium, request } from 'playwright';
const KEY = process.env.CDPFLEET_API_KEY;
const PROXY = new URL(process.env.PROXY_URL);
// Three pages, one question each: is the data in the HTML, or does it need JavaScript?
const TARGETS = [
{ url: 'https://books.toscrape.com/', item: 'product_pod', expected: 20 },
{ url: 'https://quotes.toscrape.com/js/', item: 'quote', expected: 10 },
{ url: 'https://quotes.toscrape.com/scroll', item: 'quote', expected: 10 },
];
// Step 1: a plain HTTP GET through the same proxy, no browser (Playwright's request API
// runs locally; the proxy keeps the exit IP identical to the browser's).
const http = await request.newContext({
proxy: { server: `${PROXY.protocol}//${PROXY.host}`, username: decodeURIComponent(PROXY.username), password: decodeURIComponent(PROXY.password) },
userAgent: 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/154.0.0.0 Safari/537.36',
});
const countInHtml = (html, cls) => (html.match(new RegExp(`class="[^"]*\\b${cls}\\b[^"]*"`, 'g')) || []).length;
const rows = [];
for (const t of TARGETS) {
const t0 = Date.now();
const res = await http.get(t.url, { timeout: 60000 });
const html = await res.text();
rows.push({ url: t.url, http_status: res.status(), html_kb: Math.round(html.length / 1024), items_in_html: countInHtml(html, t.item), fetch_seconds: (Date.now() - t0) / 1000 });
}
await http.dispose();
// Step 2: only the pages whose HTML didn't have the items get a browser.
const needsBrowser = rows.filter((r, i) => r.items_in_html < TARGETS[i].expected);
if (needsBrowser.length) {
const launch = await fetch('https://starter.cdpfleet.com/chromium/session', {
method: 'POST',
headers: { 'x-api-key': KEY, 'content-type': 'application/json' },
body: JSON.stringify({ proxy: process.env.PROXY_URL, headless: 'new' }),
});
if (!launch.ok) throw new Error(`launch ${launch.status} ${await launch.text()}`);
const { wsUrl } = await launch.json();
const browser = await chromium.connect(wsUrl, { headers: { 'x-api-key': KEY } });
try {
for (const r of needsBrowser) {
const t = TARGETS.find((x) => x.url === r.url);
const page = await browser.newPage();
const t0 = Date.now();
await page.goto(r.url, { timeout: 60000 });
await page.locator(`.${t.item}`).first().waitFor({ timeout: 60000 });
r.items_in_browser = await page.locator(`.${t.item}`).count();
r.browser_seconds = (Date.now() - t0) / 1000;
await page.close();
}
} finally {
await browser.close();
}
}
for (const r of rows) {
r.needs_browser = r.items_in_browser !== undefined;
r.items_in_browser ??= null;
r.browser_seconds ??= null;
}
console.log(JSON.stringify(rows, null, 2));
# pip install playwright==1.60.0 requests
# env: CDPFLEET_API_KEY, PROXY_URL
import json
import os
import re
import time
from urllib.parse import unquote, urlparse
import requests
from playwright.sync_api import sync_playwright
KEY = os.environ["CDPFLEET_API_KEY"]
PROXY = urlparse(os.environ["PROXY_URL"])
UA = "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/154.0.0.0 Safari/537.36"
# Three pages, one question each: is the data in the HTML, or does it need JavaScript?
TARGETS = [
{"url": "https://books.toscrape.com/", "item": "product_pod", "expected": 20},
{"url": "https://quotes.toscrape.com/js/", "item": "quote", "expected": 10},
{"url": "https://quotes.toscrape.com/scroll", "item": "quote", "expected": 10},
]
def count_in_html(html, cls):
return len(re.findall(rf'class="[^"]*\b{cls}\b[^"]*"', html))
with sync_playwright() as p:
# Step 1: a plain HTTP GET through the same proxy, no browser (Playwright's request API
# runs locally; the proxy keeps the exit IP identical to the browser's).
http = p.request.new_context(
proxy={"server": f"{PROXY.scheme}://{PROXY.hostname}:{PROXY.port}",
"username": unquote(PROXY.username or ""), "password": unquote(PROXY.password or "")},
user_agent=UA,
)
rows = []
for t in TARGETS:
t0 = time.time()
res = http.get(t["url"], timeout=60000)
html = res.text()
rows.append({"url": t["url"], "http_status": res.status, "html_kb": round(len(html) / 1024),
"items_in_html": count_in_html(html, t["item"]), "fetch_seconds": round(time.time() - t0, 3)})
http.dispose()
# Step 2: only the pages whose HTML didn't have the items get a browser.
needs_browser = [r for r, t in zip(rows, TARGETS) if r["items_in_html"] < t["expected"]]
if needs_browser:
launch = requests.post("https://starter.cdpfleet.com/chromium/session", headers={"x-api-key": KEY},
json={"proxy": os.environ["PROXY_URL"], "headless": "new"}, timeout=60)
if not launch.ok:
raise SystemExit(f"launch {launch.status_code} {launch.text}")
browser = p.chromium.connect(launch.json()["wsUrl"], headers={"x-api-key": KEY})
try:
for r in needs_browser:
t = next(x for x in TARGETS if x["url"] == r["url"])
page = browser.new_page()
t0 = time.time()
page.goto(r["url"], timeout=60000)
page.locator(f".{t['item']}").first.wait_for(timeout=60000)
r["items_in_browser"] = page.locator(f".{t['item']}").count()
r["browser_seconds"] = round(time.time() - t0, 3)
page.close()
finally:
browser.close()
for r in rows:
r["needs_browser"] = "items_in_browser" in r
r.setdefault("items_in_browser", None)
r.setdefault("browser_seconds", None)
print(json.dumps(rows, indent=2, ensure_ascii=False))
// Maven: com.microsoft.playwright:playwright:1.60.0, com.google.code.gson:gson:2.11.0
// Run with PLAYWRIGHT_SKIP_BROWSER_DOWNLOAD=1. env: CDPFLEET_API_KEY, PROXY_URL
import com.google.gson.*;
import com.microsoft.playwright.*;
import com.microsoft.playwright.options.Proxy;
import com.microsoft.playwright.options.RequestOptions;
import java.net.URI;
import java.net.URLDecoder;
import java.net.http.*;
import java.nio.charset.StandardCharsets;
import java.util.*;
import java.util.regex.*;
public class Main {
static final String KEY = System.getenv("CDPFLEET_API_KEY");
static final String UA = "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/154.0.0.0 Safari/537.36";
static final Gson GSON = new GsonBuilder().serializeNulls().setPrettyPrinting().disableHtmlEscaping().create();
// Three pages, one question each: is the data in the HTML, or does it need JavaScript?
record Target(String url, String item, int expected) {}
static final List<Target> TARGETS = List.of(
new Target("https://books.toscrape.com/", "product_pod", 20),
new Target("https://quotes.toscrape.com/js/", "quote", 10),
new Target("https://quotes.toscrape.com/scroll", "quote", 10));
static int countInHtml(String html, String cls) {
Matcher m = Pattern.compile("class=\"[^\"]*\\b" + Pattern.quote(cls) + "\\b[^\"]*\"").matcher(html);
int n = 0;
while (m.find()) n++;
return n;
}
public static void main(String[] args) throws Exception {
String proxyUrl = System.getenv("PROXY_URL");
URI proxy = URI.create(proxyUrl);
String[] creds = (proxy.getRawUserInfo() == null ? "" : proxy.getRawUserInfo()).split(":", 2);
String user = URLDecoder.decode(creds[0], StandardCharsets.UTF_8);
String pass = creds.length > 1 ? URLDecoder.decode(creds[1], StandardCharsets.UTF_8) : "";
try (Playwright playwright = Playwright.create()) {
// Step 1: a plain HTTP GET through the same proxy, no browser (Playwright's request API
// runs locally; the proxy keeps the exit IP identical to the browser's).
APIRequestContext http = playwright.request().newContext(new APIRequest.NewContextOptions()
.setProxy(new Proxy(proxy.getScheme() + "://" + proxy.getHost() + ":" + proxy.getPort()).setUsername(user).setPassword(pass))
.setUserAgent(UA));
List<JsonObject> rows = new ArrayList<>();
for (Target t : TARGETS) {
long t0 = System.currentTimeMillis();
APIResponse res = http.get(t.url(), RequestOptions.create().setTimeout(60000));
String html = res.text();
JsonObject row = new JsonObject();
row.addProperty("url", t.url());
row.addProperty("http_status", res.status());
row.addProperty("html_kb", Math.round(html.length() / 1024.0));
row.addProperty("items_in_html", countInHtml(html, t.item()));
row.addProperty("fetch_seconds", (System.currentTimeMillis() - t0) / 1000.0);
rows.add(row);
}
http.dispose();
// Step 2: only the pages whose HTML didn't have the items get a browser.
List<JsonObject> needsBrowser = new ArrayList<>();
for (int i = 0; i < rows.size(); i++) if (rows.get(i).get("items_in_html").getAsInt() < TARGETS.get(i).expected()) needsBrowser.add(rows.get(i));
if (!needsBrowser.isEmpty()) {
String body = "{\"proxy\": " + GSON.toJson(proxyUrl) + ", \"headless\": \"new\"}";
HttpResponse<String> launch = HttpClient.newHttpClient().send(HttpRequest.newBuilder(URI.create("https://starter.cdpfleet.com/chromium/session"))
.header("x-api-key", KEY).header("content-type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString(body)).build(), HttpResponse.BodyHandlers.ofString());
if (launch.statusCode() != 200) throw new RuntimeException("launch: " + launch.statusCode() + " " + launch.body());
String wsUrl = JsonParser.parseString(launch.body()).getAsJsonObject().get("wsUrl").getAsString();
Browser browser = playwright.chromium().connect(wsUrl, new BrowserType.ConnectOptions().setHeaders(Map.of("x-api-key", KEY)));
try {
for (JsonObject r : needsBrowser) {
Target t = TARGETS.stream().filter(x -> x.url().equals(r.get("url").getAsString())).findFirst().orElseThrow();
Page page = browser.newPage();
long t0 = System.currentTimeMillis();
page.navigate(r.get("url").getAsString(), new Page.NavigateOptions().setTimeout(60000));
page.locator("." + t.item()).first().waitFor(new Locator.WaitForOptions().setTimeout(60000));
r.addProperty("items_in_browser", page.locator("." + t.item()).count());
r.addProperty("browser_seconds", (System.currentTimeMillis() - t0) / 1000.0);
page.close();
}
} finally {
browser.close();
}
}
for (JsonObject r : rows) {
r.addProperty("needs_browser", r.has("items_in_browser"));
if (!r.has("items_in_browser")) r.add("items_in_browser", JsonNull.INSTANCE);
if (!r.has("browser_seconds")) r.add("browser_seconds", JsonNull.INSTANCE);
}
System.out.println(GSON.toJson(rows));
}
}
}
// dotnet add package Microsoft.Playwright --version 1.60.0
// env: CDPFLEET_API_KEY, PROXY_URL
using System.Diagnostics;
using System.Net.Http.Json;
using System.Text.Encodings.Web;
using System.Text.Json;
using System.Text.Json.Nodes;
using System.Text.RegularExpressions;
using Microsoft.Playwright;
const string Ua = "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/154.0.0.0 Safari/537.36";
var key = Environment.GetEnvironmentVariable("CDPFLEET_API_KEY")!;
var proxyUrl = Environment.GetEnvironmentVariable("PROXY_URL")!;
var proxy = new Uri(proxyUrl);
var creds = proxy.UserInfo.Split(':', 2);
var user = Uri.UnescapeDataString(creds[0]);
var pass = creds.Length > 1 ? Uri.UnescapeDataString(creds[1]) : "";
// Three pages, one question each: is the data in the HTML, or does it need JavaScript?
var targets = new (string Url, string Item, int Expected)[]
{
("https://books.toscrape.com/", "product_pod", 20),
("https://quotes.toscrape.com/js/", "quote", 10),
("https://quotes.toscrape.com/scroll", "quote", 10),
};
static int CountInHtml(string html, string cls) => Regex.Matches(html, $"class=\"[^\"]*\\b{Regex.Escape(cls)}\\b[^\"]*\"").Count;
using var playwright = await Playwright.CreateAsync();
// Step 1: a plain HTTP GET through the same proxy, no browser (Playwright's request API
// runs locally; the proxy keeps the exit IP identical to the browser's).
var http = await playwright.APIRequest.NewContextAsync(new()
{
Proxy = new Proxy { Server = $"{proxy.Scheme}://{proxy.Host}:{proxy.Port}", Username = user, Password = pass },
UserAgent = Ua,
});
var rows = new List<JsonObject>();
foreach (var t in targets)
{
var sw = Stopwatch.StartNew();
var res = await http.GetAsync(t.Url, new() { Timeout = 60000 });
var html = await res.TextAsync();
rows.Add(new JsonObject
{
["url"] = t.Url, ["http_status"] = res.Status, ["html_kb"] = (int)Math.Round(html.Length / 1024.0),
["items_in_html"] = CountInHtml(html, t.Item), ["fetch_seconds"] = sw.ElapsedMilliseconds / 1000.0,
});
}
await http.DisposeAsync();
// Step 2: only the pages whose HTML didn't have the items get a browser.
var needsBrowser = rows.Where((r, i) => (int)r["items_in_html"]! < targets[i].Expected).ToList();
if (needsBrowser.Count > 0)
{
using var client = new HttpClient();
client.DefaultRequestHeaders.Add("x-api-key", key);
var launch = await client.PostAsJsonAsync("https://starter.cdpfleet.com/chromium/session", new { proxy = proxyUrl, headless = "new" });
if (!launch.IsSuccessStatusCode) throw new Exception($"launch: {(int)launch.StatusCode} {await launch.Content.ReadAsStringAsync()}");
var wsUrl = (await launch.Content.ReadFromJsonAsync<JsonElement>()).GetProperty("wsUrl").GetString()!;
var browser = await playwright.Chromium.ConnectAsync(wsUrl, new() { Headers = new Dictionary<string, string> { ["x-api-key"] = key } });
try
{
foreach (var r in needsBrowser)
{
var t = targets.First(x => x.Url == (string)r["url"]!);
var page = await browser.NewPageAsync();
var sw = Stopwatch.StartNew();
await page.GotoAsync(t.Url, new() { Timeout = 60000 });
await page.Locator($".{t.Item}").First.WaitForAsync(new() { Timeout = 60000 });
r["items_in_browser"] = await page.Locator($".{t.Item}").CountAsync();
r["browser_seconds"] = sw.ElapsedMilliseconds / 1000.0;
await page.CloseAsync();
}
}
finally
{
await browser.CloseAsync();
}
}
foreach (var r in rows)
{
r["needs_browser"] = r.ContainsKey("items_in_browser");
if (!r.ContainsKey("items_in_browser")) r["items_in_browser"] = null;
if (!r.ContainsKey("browser_seconds")) r["browser_seconds"] = null;
}
Console.WriteLine(new JsonArray(rows.ToArray()).ToJsonString(new JsonSerializerOptions { WriteIndented = true, Encoder = JavaScriptEncoder.UnsafeRelaxedJsonEscaping }));
// go get github.com/playwright-community/[email protected]
// Driver: build playwright-core 1.60.0 from npm and set PLAYWRIGHT_DRIVER_PATH (see /docs/quickstart).
// env: CDPFLEET_API_KEY, PROXY_URL
package main
import (
"bytes"
"encoding/json"
"fmt"
"io"
"log"
"math"
"net/http"
"net/url"
"os"
"regexp"
"time"
"github.com/playwright-community/playwright-go"
)
var key = os.Getenv("CDPFLEET_API_KEY")
const ua = "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/154.0.0.0 Safari/537.36"
// Three pages, one question each: is the data in the HTML, or does it need JavaScript?
type target struct {
url, item string
expected int
}
var targets = []target{
{"https://books.toscrape.com/", "product_pod", 20},
{"https://quotes.toscrape.com/js/", "quote", 10},
{"https://quotes.toscrape.com/scroll", "quote", 10},
}
type row struct {
URL string `json:"url"`
HTTPStatus int `json:"http_status"`
HTMLKB int `json:"html_kb"`
ItemsInHTML int `json:"items_in_html"`
FetchSeconds float64 `json:"fetch_seconds"`
ItemsInBrowser *int `json:"items_in_browser"`
BrowserSeconds *float64 `json:"browser_seconds"`
NeedsBrowser bool `json:"needs_browser"`
}
func seconds(d time.Duration) float64 { return math.Round(d.Seconds()*1000) / 1000 }
func countInHTML(html, cls string) int {
return len(regexp.MustCompile(`class="[^"]*\b`+regexp.QuoteMeta(cls)+`\b[^"]*"`).FindAllStringIndex(html, -1))
}
func launch(name string, options map[string]any) (map[string]any, error) {
body, _ := json.Marshal(options)
req, _ := http.NewRequest("POST", "https://starter.cdpfleet.com/"+name+"/session", bytes.NewReader(body))
req.Header.Set("x-api-key", key)
req.Header.Set("content-type", "application/json")
res, err := http.DefaultClient.Do(req)
if err != nil {
return nil, err
}
defer res.Body.Close()
if res.StatusCode != http.StatusOK {
msg, _ := io.ReadAll(res.Body)
return nil, fmt.Errorf("launch %s: %s %s", name, res.Status, msg)
}
var session map[string]any
return session, json.NewDecoder(res.Body).Decode(&session)
}
func main() {
proxyURL := os.Getenv("PROXY_URL")
proxy, err := url.Parse(proxyURL)
if err != nil {
log.Fatal(err)
}
pass, _ := proxy.User.Password()
pw, err := playwright.Run(&playwright.RunOptions{SkipInstallBrowsers: true})
if err != nil {
log.Fatal(err)
}
defer pw.Stop()
// Step 1: a plain HTTP GET through the same proxy, no browser (Playwright's request API
// runs locally; the proxy keeps the exit IP identical to the browser's).
httpCtx, err := pw.Request.NewContext(playwright.APIRequestNewContextOptions{
Proxy: &playwright.Proxy{Server: proxy.Scheme + "://" + proxy.Host, Username: playwright.String(proxy.User.Username()), Password: playwright.String(pass)},
UserAgent: playwright.String(ua),
})
if err != nil {
log.Fatal(err)
}
rows := make([]*row, len(targets))
for i, t := range targets {
t0 := time.Now()
res, err := httpCtx.Get(t.url, playwright.APIRequestContextGetOptions{Timeout: playwright.Float(60000)})
if err != nil {
log.Fatal(err)
}
html, err := res.Text()
if err != nil {
log.Fatal(err)
}
rows[i] = &row{URL: t.url, HTTPStatus: res.Status(), HTMLKB: int(math.Round(float64(len(html)) / 1024)),
ItemsInHTML: countInHTML(html, t.item), FetchSeconds: seconds(time.Since(t0))}
}
httpCtx.Dispose()
// Step 2: only the pages whose HTML didn't have the items get a browser.
var needsBrowser []int
for i, r := range rows {
if r.ItemsInHTML < targets[i].expected {
needsBrowser = append(needsBrowser, i)
}
}
if len(needsBrowser) > 0 {
session, err := launch("chromium", map[string]any{"proxy": proxyURL, "headless": "new"})
if err != nil {
log.Fatal(err)
}
browser, err := pw.Chromium.Connect(session["wsUrl"].(string), playwright.BrowserTypeConnectOptions{Headers: map[string]string{"x-api-key": key}})
if err != nil {
log.Fatal(err)
}
for _, i := range needsBrowser {
t, r := targets[i], rows[i]
page, err := browser.NewPage()
if err != nil {
log.Fatal(err)
}
t0 := time.Now()
if _, err := page.Goto(t.url, playwright.PageGotoOptions{Timeout: playwright.Float(60000)}); err != nil {
log.Fatal(err)
}
if err := page.Locator("."+t.item).First().WaitFor(playwright.LocatorWaitForOptions{Timeout: playwright.Float(60000)}); err != nil {
log.Fatal(err)
}
n, err := page.Locator("." + t.item).Count()
if err != nil {
log.Fatal(err)
}
secs := seconds(time.Since(t0))
r.ItemsInBrowser, r.BrowserSeconds, r.NeedsBrowser = &n, &secs, true
page.Close()
}
browser.Close()
}
var buf bytes.Buffer
enc := json.NewEncoder(&buf)
enc.SetEscapeHTML(false)
enc.SetIndent("", " ")
enc.Encode(rows)
fmt.Print(buf.String())
}
What we got
| Page | Status | HTML (KB) | Items in HTML | Items in browser | Needed a browser | Fetch (s) | Browser (s) |
|---|---|---|---|---|---|---|---|
| https://books.toscrape.com/ | 200 | 50 | 20 | — | no | 3.55 | — |
| https://quotes.toscrape.com/js/ | 200 | 6 | 0 | 10 | yes | 1.223 | 11.043 |
| https://quotes.toscrape.com/scroll | 200 | 3 | 0 | 10 | yes | 2.605 | 6.12 |
From the Node.js run on 2026-10-05. IP addresses are replaced with placeholders (203.0.113.x); equal addresses stay equal. The other languages produced the same findings.
Raw output (Node.js)
[
{
"url": "https://books.toscrape.com/",
"http_status": 200,
"html_kb": 50,
"items_in_html": 20,
"fetch_seconds": 3.55,
"needs_browser": false,
"items_in_browser": null,
"browser_seconds": null
},
{
"url": "https://quotes.toscrape.com/js/",
"http_status": 200,
"html_kb": 6,
"items_in_html": 0,
"fetch_seconds": 1.223,
"items_in_browser": 10,
"browser_seconds": 11.043,
"needs_browser": true
},
{
"url": "https://quotes.toscrape.com/scroll",
"http_status": 200,
"html_kb": 3,
"items_in_html": 0,
"fetch_seconds": 2.605,
"items_in_browser": 10,
"browser_seconds": 6.12,
"needs_browser": true
}
]Takeaways
- One of the three pages needed no browser: the bookshop's HTML contained all 20 products, so the plain GET was the whole job — no session, no thread-seconds.
- The two JavaScript pages had zero items in their HTML (6 KB and 3 KB shells) and 10 each in the browser; the rule caught both and the browser was launched once for the pair.
- The plain fetch is the cheap probe: a second or so through the proxy, against several seconds per page for a navigation that waits for rendering — and nothing billed until the first browser launch.
- Keep the probe and the browser on the same proxy: the server then sees one visitor, not a bare HTTP client from one IP followed by a browser from another.
- Expected counts beat "is there any content": a shell page is not empty — it has a header, a footer and 6 KB of markup. Compare against what a rendered page would contain.