| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| programming:crawler [2026/08/12 09:16] – Refresh all corpus figures for the extended 2010-2026 extraction (4,322 -> 5,859 papers; crawled 859 -> 1,120). Adds a provisional 2025-2026 column; Playwright overtakes Puppeteer there. Browser fold extended (15 unmapped strings -> 1). Report script now karel.kubicek.claude | programming:crawler [2026/08/29 06:37] (current) – LLM-agent fold: new family row (4 papers), bespoke own-name 184->181, union 321->318 (28.7%->28.4%), 75->74, residue 204/210->199/205; new Agent-Driven Crawling section pointing at programming:crawler:llm_agents. Authored by Claude karel.kubicek.claude |
|---|
| Every automated web measurement makes two separate choices that papers routinely report as one: **which browser** renders the page, and **which control channel** drives it. A "Selenium crawl" says nothing about the first; a "Chrome crawl" says nothing about the second. They fail differently, they are detected differently, and they give you access to different data. | Every automated web measurement makes two separate choices that papers routinely report as one: **which browser** renders the page, and **which control channel** drives it. A "Selenium crawl" says nothing about the first; a "Chrome crawl" says nothing about the second. They fail differently, they are detected differently, and they give you access to different data. |
| |
| This page compares the options. It covers the generic automation libraries ([[#Selenium|Selenium]], [[#Puppeteer|Puppeteer]], [[#Playwright|Playwright]], [[#Plain CDP|plain CDP]]) and then the specialised privacy and security crawlers built on top of them, each of which has — or will have — its own page: [[Programming:Crawler:OpenWPM]], [[Programming:Crawler:webXray]], [[Programming:Crawler:Tracker Radar Collector]], [[Programming:Crawler:PageGraph]], [[Programming:Crawler:Foxhound]], [[Programming:Crawler:PanoptiChrome]]. | This page compares the options. It covers the generic automation libraries ([[#Selenium|Selenium]], [[#Puppeteer|Puppeteer]], [[#Playwright|Playwright]], [[#Plain CDP|plain CDP]]) and then the specialised privacy and security crawlers built on top of them, each of which has — or will have — its own page: [[Programming:Crawler:OpenWPM]], [[Programming:Crawler:webXray]], [[Programming:Crawler:Tracker Radar Collector]], [[Programming:Crawler:PageGraph]], [[Programming:Crawler:Foxhound]], [[Programming:Crawler:PanoptiChrome]]. A newer child, [[Programming:Crawler:LLM Agents]], covers the one instrument on this page that is not a control channel at all: an agent that decides the next click with a model. |
| |
| It pairs with [[Design:Automated measurements]] (whether to crawl at all), [[Design:Crawling location]] (where from), [[Programming:Stateful stateless]] (with or without a profile), [[Programming:Interaction]] (what to do on the page), and [[Privacy:Consent]] (what to do with the banner). | It pairs with [[Design:Automated measurements]] (whether to crawl at all), [[Design:Crawling location]] (where from), [[Programming:Stateful stateless]] (with or without a profile), [[Programming:Interaction]] (what to do on the page), and [[Privacy:Consent]] (what to do with the banner). |
| |
| <WRAP important> | <WRAP important> |
| The single most consequential thing on this page is not which library is best. It is that **only 11.6% of the papers in our corpus that name an automation tool also state its version** (see [[#Use in Publications]]). "We used Selenium" spans fifteen years of incompatible releases and two different wire protocols. Report the version. | The single most consequential thing on this page is not which library is best. It is that **only 12.0% of the papers in our corpus that name an automation tool also state its version** (see [[#Use in Publications]]). "We used Selenium" spans fifteen years of incompatible releases and two different wire protocols. Report the version. |
| </WRAP> | </WRAP> |
| |
| ==== Selenium ==== | ==== Selenium ==== |
| |
| Selenium is the field's default: **21.8% of all crawling papers in the corpus name it**, and it has been the most-used off-the-shelf library in every four-year window since 2014. Its virtues are real — it drives Firefox properly, it has first-class bindings in Python and Java, and code written against it in 2016 mostly still runs. | Selenium is the field's default: **21.6% of all crawling papers in the corpus name it**, and it has been the most-used off-the-shelf library in every four-year window since 2014. Its virtues are real — it drives Firefox properly, it has first-class bindings in Python and Java, and code written against it in 2016 mostly still runs. |
| |
| Its limitation was architectural. Classic WebDriver was designed to test websites, so it exposes what a user can do and hides what a debugger can see. The older escape hatch is a CDP session tunnelled over the WebDriver connection, which works but is Chromium-only and gives up the portability you chose Selenium for; the modern answer is [[#Selenium and WebDriver BiDi|WebDriver BiDi]]. This is worth stating plainly because it is the most common silent methodological compromise in the corpus: a paper says "Selenium" and reports third-party requests, and the reader cannot tell whether those came from a proxy, an extension, ''performance.getEntriesByType("resource")'', or CDP — four instruments with four different blind spots. | Its limitation was architectural. Classic WebDriver was designed to test websites, so it exposes what a user can do and hides what a debugger can see. The older escape hatch is a CDP session tunnelled over the WebDriver connection, which works but is Chromium-only and gives up the portability you chose Selenium for; the modern answer is [[#Selenium and WebDriver BiDi|WebDriver BiDi]]. This is worth stating plainly because it is the most common silent methodological compromise in the corpus: a paper says "Selenium" and reports third-party requests, and the reader cannot tell whether those came from a proxy, an extension, ''performance.getEntriesByType("resource")'', or CDP — four instruments with four different blind spots. |
| ==== Puppeteer ==== | ==== Puppeteer ==== |
| |
| Puppeteer is a thin, well-documented wrapper over CDP maintained by the Chrome team. It appears in **6.3% of crawling papers**, first in 2018, and it is what most of the modern research tooling is built on — DuckDuckGo's [[Programming:Crawler:Tracker Radar Collector|Tracker Radar Collector]] and Brave's ''pagegraph-crawl'' are both Puppeteer programs. Recent versions download and pin a //Chrome for Testing// build, which fixes the "which Chrome was that?" reproducibility problem for free, provided you record the version. | Puppeteer is a thin, well-documented wrapper over CDP maintained by the Chrome team. It appears in **6.8% of crawling papers**, first in 2018, and it is what most of the modern research tooling is built on — DuckDuckGo's [[Programming:Crawler:Tracker Radar Collector|Tracker Radar Collector]] and Brave's ''pagegraph-crawl'' are both Puppeteer programs. Recent versions download and pin a //Chrome for Testing// build, which fixes the "which Chrome was that?" reproducibility problem for free, provided you record the version. |
| |
| Puppeteer used to be Chromium-and-CDP all the way down. It is not any more: it supports Firefox over WebDriver BiDi, and when you launch Firefox, BiDi is the default. Chrome still defaults to CDP "since not all CDP features are supported by WebDriver BiDi yet", and a Puppeteer feature that BiDi cannot express throws ''UnsupportedOperation''.((Puppeteer documentation, [[https://pptr.dev/webdriver-bidi|WebDriver BiDi support]], checked 2026-08-06.)) So Puppeteer is now a plausible cross-engine choice, with the caveat that the two engines do not give you the same API surface — check that the calls your measurement depends on work on both before you report a cross-browser comparison. | Puppeteer used to be Chromium-and-CDP all the way down. It is not any more: it supports Firefox over WebDriver BiDi, and when you launch Firefox, BiDi is the default. Chrome still defaults to CDP "since not all CDP features are supported by WebDriver BiDi yet", and a Puppeteer feature that BiDi cannot express throws ''UnsupportedOperation''.((Puppeteer documentation, [[https://pptr.dev/webdriver-bidi|WebDriver BiDi support]], checked 2026-08-06.)) So Puppeteer is now a plausible cross-engine choice, with the caveat that the two engines do not give you the same API surface — check that the calls your measurement depends on work on both before you report a cross-browser comparison. |
| Playwright is the newest of the three and the only one that drives three different engines — Chromium, Firefox and WebKit — through **browser builds it ships itself**, pinned per Playwright version. For a measurement study that is the interesting property: the browser binary is part of the dependency, so naming the Playwright version pins the engine too, and a cross-engine comparison stops requiring three separate harnesses. | Playwright is the newest of the three and the only one that drives three different engines — Chromium, Firefox and WebKit — through **browser builds it ships itself**, pinned per Playwright version. For a measurement study that is the interesting property: the browser binary is part of the dependency, so naming the Playwright version pins the engine too, and a cross-engine comparison stops requiring three separate harnesses. |
| |
| It is still rare in the literature — **10 papers, all of them 2022 or later**, the youngest tool on this page — but it is the only one whose adoption curve is still going up steeply, and the small group using it reports its version far more often than anyone else (see [[#Use in Publications]]). | It is the youngest tool on this page — **34 papers (3.0% of crawling papers), all of them 2022 or later** — and the only one whose adoption curve is still going up steeply: in the provisional 2025–2026 window it **overtakes Puppeteer**, 9.6% against 8.1%. Its users also report its version better than most, though the margin is narrow: 26.5%, just ahead of OpenWPM's 25.9% and more than double everyone else (see [[#Use in Publications]]). |
| |
| Three caveats for measurement work. | Three caveats for measurement work. |
| ==== Plain CDP ==== | ==== Plain CDP ==== |
| |
| You can skip the libraries entirely: launch Chromium with ''--remote-debugging-port'', open a WebSocket, enable the domains you want, and read the events. **35 papers (4.1%) did**, most of them between 2018 and 2021 — a window that closes as Puppeteer matures. | You can skip the libraries entirely: launch Chromium with ''--remote-debugging-port'', open a WebSocket, enable the domains you want, and read the events. **42 papers (3.8%) did**, most of them between 2018 and 2021 — a window that closes as Puppeteer matures. |
| |
| Reach for it when your analysis lives in a language with no good automation binding, when you need a CDP domain the wrapper does not expose, or when you want to know exactly what your instrument is doing. The cost, visible in the code below, is that everything a library gives you is now yours to write: waiting for the port, waiting for the load, waiting for the network to settle, and reconnecting when a target dies. | Reach for it when your analysis lives in a language with no good automation binding, when you need a CDP domain the wrapper does not expose, or when you want to know exactly what your instrument is doing. The cost, visible in the code below, is that everything a library gives you is now yours to write: waiting for the port, waiting for the load, waiting for the network to settle, and reconnecting when a target dies. |
| | VisibleV8 {[jueckstock2019_visiblev8]} | patched Chromium (V8) | any CDP library | native-code JS API tracing inside V8, below the reach of page JavaScript | append-only trace logs | custom Chromium build | **9** | | | VisibleV8 {[jueckstock2019_visiblev8]} | patched Chromium (V8) | any CDP library | native-code JS API tracing inside V8, below the reach of page JavaScript | append-only trace logs | custom Chromium build | **9** | |
| | [[Programming:Crawler:Foxhound|SAP Project Foxhound]] | patched Firefox (Gecko + SpiderMonkey) | Playwright((The only automation layer the repository documents; it carries its own ''.PLAYWRIGHT_VERSION'' and a CI workflow for it.)) | **dynamic taint tracking**: string-level flows from sources to sinks | ''__taintreport'' DOM events | full Firefox build toolchain, or prebuilt binaries | **8** | | | [[Programming:Crawler:Foxhound|SAP Project Foxhound]] | patched Firefox (Gecko + SpiderMonkey) | Playwright((The only automation layer the repository documents; it carries its own ''.PLAYWRIGHT_VERSION'' and a CI workflow for it.)) | **dynamic taint tracking**: string-level flows from sources to sinks | ''__taintreport'' DOM events | full Firefox build toolchain, or prebuilt binaries | **8** | |
| | [[Programming:Crawler:PanoptiChrome|PanoptiChrome]] {[kanyal2024_panoptichrome]} | patched Chromium/V8, pinned to Chrome 116 | — | dynamic taint tracking in Chromium, arbitrary sources and sinks | //see the paper// | build a pinned Chrome 116 fork from patches | **2** | | | [[Programming:Crawler:PanoptiChrome|PanoptiChrome]] {[kanyal2024_panoptichrome]} | patched Chromium/V8, pinned to Chrome 116 | — | dynamic taint tracking in Chromium, on objects rather than strings, including implicit flows | one plain-text log per V8 isolate, in VisibleV8's format | build a pinned Chrome 116 fork from patches — the patch published with the paper **does not compile**; use the completed one, or a third party's prebuilt binary | **2** | |
| | [[Programming:Crawler:webXray|webXray]] {[libert2015_invisible]} | PhantomJS //(historically)// | — | third-party request detection plus **domain-to-company attribution** | MySQL, CSV reports | //see below// | **7** | | | [[Programming:Crawler:webXray|webXray]] {[libert2015_invisible]} | PhantomJS in 2015; **consumer Chrome over raw CDP** in the last public version | — | third-party request detection plus **domain-to-company attribution** | SQLite or Postgres, CSV reports | //see below// | **7** | |
| |
| Notes that matter more than the table: | Notes that matter more than the table: |
| * **PageGraph does not track your automation.** Its own documentation warns that "PageGraph currently does not track puppeteer / automation scripts, and so modifying or interacting with the document through devtools/puppeteer while recording a PageGraph file will likely fail".((''pagegraph-crawl'' README, [[https://github.com/brave/pagegraph-crawl|github.com/brave/pagegraph-crawl]], checked 2026-08-06.)) It is a recorder of what the page did, not a driver of interaction. Its own limitations list is unusually honest and worth reading before you rely on it: no WebSocket tracking, no worker tracking, request headers not recorded (responses only), no ''style='' mutation tracking.((PageGraph wiki, [[https://github.com/brave/brave-browser/wiki/PageGraph|"Would Be Nice / Someday / Known Limitations"]], checked 2026-08-06. Query tooling: ''pagegraph-query'' (Python) under ''brave-experiments'', ''pagegraph-rust'' under ''brave''.)) | * **PageGraph does not track your automation.** Its own documentation warns that "PageGraph currently does not track puppeteer / automation scripts, and so modifying or interacting with the document through devtools/puppeteer while recording a PageGraph file will likely fail".((''pagegraph-crawl'' README, [[https://github.com/brave/pagegraph-crawl|github.com/brave/pagegraph-crawl]], checked 2026-08-06.)) It is a recorder of what the page did, not a driver of interaction. Its own limitations list is unusually honest and worth reading before you rely on it: no WebSocket tracking, no worker tracking, request headers not recorded (responses only), no ''style='' mutation tracking.((PageGraph wiki, [[https://github.com/brave/brave-browser/wiki/PageGraph|"Would Be Nice / Someday / Known Limitations"]], checked 2026-08-06. Query tooling: ''pagegraph-query'' (Python) under ''brave-experiments'', ''pagegraph-rust'' under ''brave''.)) |
| * **Tracker Radar Collector caps concurrency at 38** crawlers, injects a simple anti-bot-detection script into every frame by default, and can be pointed at a Selenium Hub or a specific Chromium version.((''tracker-radar-collector'' README, [[https://github.com/duckduckgo/tracker-radar-collector|github.com/duckduckgo/tracker-radar-collector]], checked 2026-08-06.)) It has no canonical academic paper; cite the repository and the Tracker Radar dataset. | * **Tracker Radar Collector caps concurrency at 38** crawlers, injects a simple anti-bot-detection script into every frame by default, and can be pointed at a Selenium Hub or a specific Chromium version.((''tracker-radar-collector'' README, [[https://github.com/duckduckgo/tracker-radar-collector|github.com/duckduckgo/tracker-radar-collector]], checked 2026-08-06.)) It has no canonical academic paper; cite the repository and the Tracker Radar dataset. |
| * **webXray needs care.** The repository its own documentation points at, ''github.com/timlib/webXray'', **no longer exists** — the account has no public repositories at all, and there is no ''webxray'' package on PyPI. What survives on GitHub are stale third-party mirrors, the newest of which was last pushed in 2015 and targets PhantomJS, whose own maintainer suspended development in 2018.(([[https://github.com/ariya/phantomjs/issues/15344|ariya/phantomjs#15344]], "Archiving the project: suspending the development", opened 2018-03-03.)) ''webxray.org'' is a live landing page with no source link.((Checked 2026-08-06: ''api.github.com/repos/timlib/webXray'' → HTTP 404; ''api.github.com/users/timlib/repos'' → empty; ''pypi.org/pypi/webxray/json'' → not found.)) Treat the 2015 paper {[libert2015_invisible]} as the citation for the //method// — third-party request measurement with company attribution — and not as a tool you can install today. <wrap todo>If you know where webXray is currently developed, please correct this.</wrap> | * **webXray cannot be installed, and the reason is now known.** The repository its own documentation points at, ''github.com/timlib/webXray'', **no longer exists**, the ''timlib'' account has no public repositories, and there is no ''webxray'' package on PyPI. webXray is now a **commercial** product, ''webxray.ai'', run by webXray LLC with Libert as founder and CEO; ''webxray.org'' has been reduced to a placeholder.((Checked 2026-08-17: ''api.github.com/repos/timlib/webXray'' → HTTP 404, ''users/timlib'' → HTTP 200 with ''public_repos: 0'' and ''company: webXray.ai''; ''pypi.org/pypi/webxray/json'' → 404; ''https://timlibert.me/'' states "(Dr.) Timothy Libert is founder and CEO of webXray LLC".)) The most complete surviving copy is ''thezedwards/webXray'', last pushed **2021-03-04** — webXray 3.x, which drives **consumer Chrome over the DevTools protocol**, not PhantomJS; its licence is PolyForm Strict 1.0.0, which permits noncommercial use but **forbids redistribution**, so there is no lawful route to the code now that upstream is gone. Treat {[libert2015_invisible]} and {[libert2018_automated]} as citations for the //method// — third-party request measurement with company attribution — and see [[Programming:Crawler:webXray]] for what survives of it: the ownership database, and how it compares to Tracker Radar and Disconnect today. |
| * **Taint tracking is a different question.** Foxhound and PanoptiChrome answer "did this value reach that sink?", not "what did the page load". If your question is about tracking prevalence, they are the wrong instrument and cost you a browser build; if it is about how data escapes — client-side XSS, DOM-based leaks, fingerprinting inputs — nothing else answers it. Foxhound tells you what to cite: its README's "Cite us!" section asks for the EuroS&P paper in which the browser is described, Klein et al. {[klein2022_handsanitizers]}, and its wiki separately lists the papers that have used it. | * **Taint tracking is a different question.** Foxhound and PanoptiChrome answer "did this value reach that sink?", not "what did the page load". If your question is about tracking prevalence, they are the wrong instrument and cost you a browser build; if it is about how data escapes — client-side XSS, DOM-based leaks, fingerprinting inputs — nothing else answers it. **Between the two, start with Foxhound**: the one head-to-head evaluation measured PanoptiChrome at 50% compatibility and 36.7× overhead against Foxhound's 95% and 1.4×, and no paper in this corpus has used PanoptiChrome as an instrument {[calzavara2025_dynamic]} — see [[Programming:Crawler:PanoptiChrome]] for when it is nevertheless the only option. Foxhound tells you what to cite: its README's "Cite us!" section asks for the EuroS&P paper in which the browser is described, Klein et al. {[klein2022_handsanitizers]}, and its wiki separately lists the papers that have used it. |
| |
| ===== Being Detected ===== | ===== Agent-Driven Crawling ===== |
| |
| Whatever you drive, the website may notice. This is a measurement-validity problem, not just an engineering nuisance: a bot-managed site serves your crawler a different page, and that difference is silently attributed to whatever you were studying. | Since 2025 there is a fifth option that does not fit the two-layer model above: a |
| | **crawler whose next action is chosen by a model** rather than scripted. The open-source |
| | ones — Browser Use, BrowserGym/AgentLab, Skyvern — still drive Chromium through CDP or |
| | Playwright, so they add no new control channel; the vendor computer-use APIs act on |
| | screenshots and coordinates and leave the channel to you. What all of them change is that |
| | the sequence of actions is non-deterministic, the completion rate becomes an empirical |
| | quantity, and the cost per site is measured in cents rather than fractions of one. |
| | |
| | **Five of the 1,120 crawling papers here drove a crawl this way, all of them in 2026.**((Four |
| | in the table below, which counts only ''tools[]'' entries in an automation category, as |
| | every other row of that table does. A fifth paper drove its crawl with Claude's Computer |
| | Use API, which the extraction filed under category ''llm''; it is counted on the child |
| | page and explained there.)) That is not enough to call it practice, and this page does not |
| | recommend it as a default. It is enough for a page of its own, because those five — and |
| | four neighbours that measure agents rather than crawl with them — have already published |
| | completion rates, per-site costs and (on benchmark tasks) run-to-run variance that nobody |
| | starting an agent crawl should have to rediscover. |
| | **[[Programming:Crawler:LLM Agents]]** has those numbers, the tool currency, and what to |
| | report. |
| | |
| | ===== Being Detected ===== |
| |
| * **Automation is detectable in the browser.** Vastel et al. {[vastel2018_scanner]} show that fingerprint inconsistencies distinguish instrumented and spoofed browsers from ordinary ones; headless Chrome and driver-injected properties are among the easiest signals to read. | Whatever you drive, the website may notice, and a bot-managed site serves your crawler a |
| * **Specific frameworks are detectable specifically.** Krumnow et al. {[krumnow2022_gullible]} study OpenWPM detection in the wild. A tool that 58 papers in this corpus share is a tool worth writing a detector for. | different page — a measurement-validity problem, not an engineering nuisance. |
| * **Crawls differ from humans even when nobody is detecting you.** Zeber et al. {[zeber2020representativeness]} quantified how far automated crawls diverge from real browsing on common tracking and fingerprinting metrics, and found crawls fail to capture the diversity of user environments. | **[[Programming:Crawler Detection]] covers it in full**: what each layer of the stack |
| * **The vantage point compounds this**, since datacenter IP ranges are themselves a low-trust signal {[jueckstock2021_realistic]} — see [[Design:Crawling location]]. | gives you away with, why cloaking is the shape that ends up in your abstract, how to |
| * **Repeating a crawl is not free either.** Demir et al. {[demir2022_reproducibility]} found substantial variation between repetitions of the same web measurement, so a single crawl of a single configuration is a point estimate with unstated error bars. | measure the obstruction instead of fighting it, and what the corpus says about how rarely |
| | anyone reports it. Two points from there bear on the choice made //on this page//: |
| |
| Nine papers in the corpus reach for explicit anti-detection patches (''puppeteer-extra-plugin-stealth'', ''undetected-chromedriver''), all of them 2021 or later. Two warnings if you follow them. They are an arms race you will lose quietly and without notice, so any result that depends on them needs a validity check that does not. And evading bot management is a decision with an ethics dimension — see [[Practices:Ethics]] — because you are deliberately overriding a site operator's expressed access preference. | * **The framework is detectable, not just automation in general.** Vastel et al. {[vastel2018_scanner]} show fingerprint inconsistencies distinguish instrumented browsers from ordinary ones, and Krumnow et al. {[krumnow2022_gullible]} show OpenWPM specifically has detectors deployed against it in the wild. |
| | * **The anti-detection patches in the tables below are an arms race you lose quietly**, and most of those packages are now unmaintained or have moved to a successor under a different name; [[Programming:Crawler Detection]] dates each of them. Using them is also an ethics decision — see [[Practices:Ethics]] — because you are overriding a site operator's expressed access preference. |
| |
| ===== Use in Publications ===== | ===== Use in Publications ===== |
| ^ Family ^ Papers ^ Share of 1,120 ^ | ^ Family ^ Papers ^ Share of 1,120 ^ |
| | Selenium | 242 | 21.6% | | | Selenium | 242 | 21.6% | |
| | //Bespoke crawler, given its own name// (''SSOScan'', ''AdFisher'', ''PhishPrint'', ''CryptoScamTracker''…) | 184 | 16.4% | | | //Bespoke crawler, given its own name// (''SSOScan'', ''AdFisher'', ''PhishPrint'', ''CryptoScamTracker''…) | 181 | 16.2% | |
| | //Bespoke crawler, described generically// ("our crawler", "a custom Python crawler", "a crawling extension") | 147 | 13.1% | | | //Bespoke crawler, described generically// ("our crawler", "a custom Python crawler", "a crawling extension") | 147 | 13.1% | |
| | Puppeteer | 76 | 6.8% | | | Puppeteer | 76 | 6.8% | |
| | Tor Browser Crawler / ''tbselenium'' | 20 | 1.8% | | | Tor Browser Crawler / ''tbselenium'' | 20 | 1.8% | |
| | App-store and platform scrapers | 18 | 1.6% | | | App-store and platform scrapers | 18 | 1.6% | |
| | Third-party browser extension used as the instrument (Ghostery, uBlock Origin, NoScript…)((An extension the authors wrote themselves is a bespoke crawler and is counted in the two bespoke rows instead, because that is the distinction that decides whether a reader can identify the instrument.)) | 14 | 1.3% | | |
| | HTTP-level scrapers, no browser (BeautifulSoup, ''requests'') | 17 | 1.5% | | | HTTP-level scrapers, no browser (BeautifulSoup, ''requests'') | 17 | 1.5% | |
| | Web-performance harnesses (WebPageTest, Browsertime, Lighthouse) | 16 | 1.4% | | | Web-performance harnesses (WebPageTest, Browsertime, Lighthouse) | 16 | 1.4% | |
| | | Third-party browser extension used as the instrument (Ghostery, uBlock Origin, NoScript…)((An extension the authors wrote themselves is a bespoke crawler and is counted in the two bespoke rows instead, because that is the distinction that decides whether a reader can identify the instrument.)) | 14 | 1.3% | |
| | OS-level GUI automation (''pyautogui'', ''xdotool'') | 14 | 1.3% | | | OS-level GUI automation (''pyautogui'', ''xdotool'') | 14 | 1.3% | |
| | Archival crawlers (Heritrix, HTTrack, ''wget -r'') | 12 | 1.1% | | | Archival crawlers (Heritrix, HTTrack, ''wget -r'') | 12 | 1.1% | |
| | Anti-detection patches (''stealth'', ''undetected-chromedriver'') | 9 | 0.8% | | | Anti-detection patches (''stealth'', ''undetected-chromedriver'') | 9 | 0.8% | |
| | Fuzzers and monkey testers | 5 | 0.4% | | | Fuzzers and monkey testers | 5 | 0.4% | |
| | | [[Programming:Crawler:LLM Agents|LLM browser agents]] (Browser Use, BrowserGym/AgentLab)((Added as a family on 2026-08-29. All four papers are 2026, i.e. entirely inside the provisional slice. The family is tested //before// Playwright and Puppeteer because several of these agents are built on those libraries, so a name matched in the wrong order would land in the library's row instead. A fifth crawling paper drove its crawl with an agent the extraction filed under category ''llm'', outside this table's population; [[Programming:Crawler:LLM Agents]] counts it and explains the boundary.)) | 4 | 0.4% | |
| | webXray | 1 | 0.1% | | | webXray | 1 | 0.1% | |
| |
| <WRAP important> | <WRAP important> |
| Put the two bespoke rows together — 10 papers are in both — and **321 of 1,120 crawling papers (28.7%) crawled with something home-grown or too obscure to have a family here.** That is more than used Selenium (242), and more than the number naming any of Puppeteer, Playwright, direct CDP, Scrapy or OpenWPM (215). Of the 184 that gave their crawler a proper name, **75 named nothing else at all** — the name is the only identification of the instrument in the paper, and it means nothing to a reader who does not have the code. | Put the two bespoke rows together — 10 papers are in both — and **318 of 1,120 crawling papers (28.4%) crawled with something home-grown or too obscure to have a family here.** That is more than used Selenium (242), and more than the number naming any of Puppeteer, Playwright, direct CDP, Scrapy or OpenWPM (215). Of the 181 that gave their crawler a proper name, **74 named nothing else at all** — the name is the only identification of the instrument in the paper, and it means nothing to a reader who does not have the code. |
| </WRAP> | </WRAP> |
| |
| | **Any automation tool** | **723** | **87** | **12.0%** | | | **Any automation tool** | **723** | **87** | **12.0%** | |
| |
| This is the largest reporting gap on the page and the cheapest to fix. OpenWPM users are more than twice as good as average at it, plausibly because OpenWPM's own releases are numbered and cited; Playwright's cohort is next best. Everyone else writes "we used Selenium". | This is the largest reporting gap on the page and the cheapest to fix. Playwright's cohort is best at it (26.5%) and OpenWPM's is a fraction behind (25.9%) — both more than double the 12.0% average, and in OpenWPM's case plausibly because its own releases are numbered and cited. On the previous corpus Playwright led by a wide margin on a base of ten papers; with 34 the two are effectively tied. Everyone else writes "we used Selenium". |
| |
| ==== Configuration reporting, and whether the tool predicts it ==== | ==== Configuration reporting, and whether the tool predicts it ==== |
| | ''tbselenium'' / tor-browser-crawler | 21 | 0 | 2017–2026 | | | ''tbselenium'' / tor-browser-crawler | 21 | 0 | 2017–2026 | |
| | Tracker Radar Collector | 21 | 1 | 2021–2026 | | | Tracker Radar Collector | 21 | 1 | 2021–2026 | |
| | Anti-detection patches | 12 | 2 | 2021–2026 | | | Anti-detection patches | 11 | 1 | 2021–2026 | |
| | VisibleV8 | 9 | 0 | 2019–2026 | | | VisibleV8 | 9 | 0 | 2019–2026 | |
| | Brave PageGraph | 8 | 0 | 2020–2025 | | | Brave PageGraph | 8 | 0 | 2020–2025 | |
| |
| * **How they were produced.** One structured record per paper was extracted from full text, each tuple carrying a verbatim evidence quote and its section, so any figure here can be traced to the sentence that supports it. The script that produces every number on this page, with its denominators, is ''report_crawler.mjs''; the folding rules are in ''tool_fold.mjs''. Every query, the script's unedited output and the full residue are on [[provenance:programming:crawler]]; corpus-level caveats are on [[literature:corpus]]. | * **How they were produced.** One structured record per paper was extracted from full text, each tuple carrying a verbatim evidence quote and its section, so any figure here can be traced to the sentence that supports it. The script that produces every number on this page, with its denominators, is ''report_crawler.mjs''; the folding rules are in ''tool_fold.mjs''. Every query, the script's unedited output and the full residue are on [[provenance:programming:crawler]]; corpus-level caveats are on [[literature:corpus]]. |
| * **How names were folded.** Tool names are free text and agree run-to-run on only about a fifth of exact strings, so nothing here is counted by exact string. Names were folded into the families shown by an explicit, ordered list of regular expressions — specific tools before the generic libraries they are built on, so ''puppeteer-extra-plugin-stealth'' lands in //Anti-detection patches// and not in //Puppeteer//. Folding matters: the exact string ''Selenium'' appears in 194 papers, while the folded family covers 242, so counting exact strings would undercount Selenium by 19.8% — the difference is ''Selenium WebDriver'', ''Selenium Webdriver'', ''Python Selenium WebDriver'', ''selenium'', ''ChromeDriver'' and ''Selenium's ChromeDriver''. Of **1,075 tool mentions across 501 distinct strings** in the crawling population, **204 distinct strings across 210 mentions match no family**. They are not discarded — they are the //Bespoke crawler, given its own name// row, because almost all of them are one paper's own tool (''SSOScan'', ''AdFisher'', ''Formlock'', ''CryptoScamTracker'', ''Spider-Scents'': 184 papers, essentially one name each). A handful are third-party tools we chose not to give a family of their own (''OmniCrawl'', ''JAW'', ''BrowserStack'', ''MetaMask automator'', ''Headless Chromium'', the ''measurement framework of Demir et al.''), so read that row as "home-grown or obscure" rather than strictly "home-grown". ''report_crawler.mjs'' prints the full list, so nothing vanishes. Browser names folded to 1 unclassified string out of 529 papers (''Ghostery'', which is an extension, not a browser). | * **How names were folded.** Tool names are free text and agree run-to-run on only about a fifth of exact strings, so nothing here is counted by exact string. Names were folded into the families shown by an explicit, ordered list of regular expressions — specific tools before the generic libraries they are built on, so ''puppeteer-extra-plugin-stealth'' lands in //Anti-detection patches// and not in //Puppeteer//. Folding matters: the exact string ''Selenium'' appears in 194 papers, while the folded family covers 242, so counting exact strings would undercount Selenium by 19.8% — the difference is ''Selenium WebDriver'', ''Selenium Webdriver'', ''Python Selenium WebDriver'', ''selenium'', ''ChromeDriver'' and ''Selenium's ChromeDriver''. Of **1,075 tool mentions across 501 distinct strings** in the crawling population, **199 distinct strings across 205 mentions match no family**. They are not discarded — they are the //Bespoke crawler, given its own name// row, because almost all of them are one paper's own tool (''SSOScan'', ''AdFisher'', ''Formlock'', ''CryptoScamTracker'', ''Spider-Scents'': 181 papers, essentially one name each). A handful are third-party tools we chose not to give a family of their own (''OmniCrawl'', ''JAW'', ''BrowserStack'', ''MetaMask automator'', ''Headless Chromium'', the ''measurement framework of Demir et al.''), so read that row as "home-grown or obscure" rather than strictly "home-grown". ''report_crawler.mjs'' prints the full list, so nothing vanishes. Browser names folded to 1 unclassified string out of 529 papers (''Ghostery'', which is an extension, not a browser). |
| * **A paper counts once per family**, never once per mention, and shares do not sum to 100% because a paper can name several tools. A family's share is of the population named in its heading. | * **A paper counts once per family**, never once per mention, and shares do not sum to 100% because a paper can name several tools. A family's share is of the population named in its heading. |
| * **Silence is not absence.** "Does not name a framework" means the paper did not say, not that the authors used none. These are reporting figures. | * **Silence is not absence.** "Does not name a framework" means the paper did not say, not that the authors used none. These are reporting figures. |
| * **''used'' and ''produced'' both count** as driving a crawl — a paper that built its own crawler crawled with it. Tools only //compared// or //cited// do not, which is why the last table separates the two. | * **''used'' and ''produced'' both count** as driving a crawl — a paper that built its own crawler crawled with it. Tools only //compared// or //cited// do not, which is why the last table separates the two. |
| * **Quotes were spot-checked**, by ''scripts/quote_check.mjs'' rather than by hand. Of the **106** evidence quotes attached to an OpenWPM or Playwright tool mention anywhere in the corpus, 58 match the source text exactly once whitespace and line-break hyphenation are normalised, and 32 more match on at least 60% of their five-word windows. The remaining 16 were then read by hand against the full text: **all sixteen are present in the paper and none was unsupported.** The mismatch is always either a bracketed citation marker the extraction dropped (''the Open-WPM platform [52] on the Firefox browser'') or the two-column reading order still interleaving mid-sentence (''we created a Playwright**constrains the valid child elements, and everything else is moved based** [18] crawler''). | * **Quotes were spot-checked**, by ''scripts/quote_check.mjs'' rather than by hand. Of the **106** evidence quotes attached to an OpenWPM or Playwright tool mention anywhere in the corpus, 58 match the source text exactly once whitespace and line-break hyphenation are normalised, and 32 more match on at least 60% of their five-word windows. The remaining 16 were then read by hand against the full text: **all sixteen are present in the paper and none was unsupported.** The mismatch is always either a bracketed citation marker the extraction dropped (''the Open-WPM platform [52] on the Firefox browser'') or the two-column reading order still interleaving mid-sentence (''we created a Playwright**constrains the valid child elements, and everything else is moved based** [18] crawler''). |
| * **Venue coverage.** EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS are absent entirely; any claim here is a claim about seven venues. Notably, several of the tool papers this page recommends were published //outside// them — Krumnow et al. at CoNEXT, Klein et al. at EuroS&P, Libert in a communication journal — so a tool's paper count in this table is a lower bound on its standing. | * **Venue coverage.** Seven venues only; the scope, the selection funnel and what each stage of it costs are on [[literature:corpus]]. Notably, several of the tool papers this page recommends were published //outside// them — Krumnow et al. at CoNEXT, Klein et al. at EuroS&P, Libert in a communication journal — so a tool's paper count in this table is a lower bound on its standing. |
| * **''crawled'' is defined** as a paper whose crawl configuration was recorded or whose study types include an automated web crawl (1,120 papers, 19.1% of the corpus). This corpus is seven broad security venues, not a web-measurement corpus, so shares of all 5,859 papers would be meaningless here. | * **''crawled'' is defined** as a paper whose crawl configuration was recorded or whose study types include an automated web crawl (1,120 papers, 19.1% of the corpus). This corpus is seven broad security venues, not a web-measurement corpus, so shares of all 5,859 papers would be meaningless here. |
| |
| ===== Recommendations ===== | ===== Recommendations ===== |
| |
| - **Report both layers and both versions.** "Chromium 121 driven by Playwright 1.41" is a complete answer; "Selenium" is not. This costs one sentence and is the difference between a replicable and an unreplicable crawl. It is also the field's biggest gap: 11.6%. | - **Report both layers and both versions.** "Chromium 121 driven by Playwright 1.41" is a complete answer; "Selenium" is not. This costs one sentence and is the difference between a replicable and an unreplicable crawl. It is also the field's biggest gap: 12.0%. |
| - **Pick the instrument from the question, not the habit.** | - **Pick the instrument from the question, not the habit.** |
| * Third-party requests, cookies, response bodies → Puppeteer, Playwright or CDP. With Selenium, use WebDriver BiDi and verify what your binding actually delivers; never classic WebDriver alone. | * Third-party requests, cookies, response bodies → Puppeteer, Playwright or CDP. With Selenium, use WebDriver BiDi and verify what your binding actually delivers; never classic WebDriver alone. |
| * Multi-browser or multi-language, or a lab already fluent in it → Selenium. | * Multi-browser or multi-language, or a lab already fluent in it → Selenium. |
| * The lowest-effort modern tracking crawl → Tracker Radar Collector. | * The lowest-effort modern tracking crawl → Tracker Radar Collector. |
| - **Do not write a new crawler for a solved problem.** A quarter of crawling papers used a home-grown one, and 34 of them identify it by nothing but a name they invented. If you must build one, say what it is built on, and publish it. | - **Do not write a new crawler for a solved problem.** Nearly three in ten crawling papers used a home-grown one, and 75 of them identify it by nothing but a name they invented. If you must build one, say what it is built on, and publish it. |
| - **State headless or headful and justify it.** Headless is the most detectable configuration you can choose {[vastel2018_scanner]} and only 12.9% of papers say which they used. | - **State headless or headful and justify it.** Headless is the most detectable configuration you can choose {[vastel2018_scanner]} and only 12.5% of papers say which they used. |
| - **Pin the browser binary, not just the library.** Puppeteer and Playwright do this for you; Selenium does not. Record the exact build and archive it with your research artefact (see [[:Artifacts]], a page this wiki still owes you). | - **Pin the browser binary, not just the library.** Puppeteer and Playwright do this for you; Selenium does not. Record the exact build and archive it with your research artefact (see [[:Artifacts]], a page this wiki still owes you). |
| - **Validate against something that is not your crawler.** A manual visit to a sample, a second browser engine, or a second vantage point {[jueckstock2021_realistic]}. Crawls diverge from human browsing in ways your crawl cannot see {[zeber2020representativeness]}, and repeated crawls diverge from each other {[demir2022_reproducibility]}. | - **Validate against something that is not your crawler.** A manual visit to a sample, a second browser engine, or a second vantage point {[jueckstock2021_realistic]}. Crawls diverge from human browsing in ways your crawl cannot see {[zeber2020representativeness]}, and repeated crawls diverge from each other {[demir2022_reproducibility]}. |
| ===== Open Questions ===== | ===== Open Questions ===== |
| |
| <wrap todo> | <WRAP todo> |
| * A head-to-head comparison of what OpenWPM, Tracker Radar Collector and a plain Playwright crawl each detect on the same sample, with the same vantage point and the same date, does not exist in the literature we found. It would settle a design question the whole field guesses at. | * A head-to-head comparison of what OpenWPM, Tracker Radar Collector and a plain Playwright crawl each detect on the same sample, with the same vantage point and the same date, does not exist in the literature we found. It would settle a design question the whole field guesses at. |
| * Playwright's patched Firefox and WebKit builds are used as stand-ins for the real browsers. How far the patching moves the fingerprint, and whether it changes what trackers do, is unmeasured. | * Playwright's patched Firefox and WebKit builds are used as stand-ins for the real browsers. How far the patching moves the fingerprint, and whether it changes what trackers do, is unmeasured. |
| * Where webXray is developed today (see [[#Specialised Measurement Crawlers|above]]). | |
| * Whether Selenium's steady quarter-share reflects a considered choice or institutional inertia — a question for a survey of authors, not for this corpus. | * Whether Selenium's steady quarter-share reflects a considered choice or institutional inertia — a question for a survey of authors, not for this corpus. |
| </wrap> | </WRAP> |
| |
| ====== References ====== | ====== References ====== |