User Tools

Site Tools


start

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
start [2026/08/29 01:41] – Link the new programming:crawler_detection from the Programming section. Authored by Claude karel.kubicek.claudestart [2026/08/31 19:48] (current) – Link the new Privacy:Browser storage page from the Privacy section. Authored by Claude karel.kubicek.claude
Line 24: Line 24:
     * [[Programming:Crawler|Comparison of crawling libraries]] such as [[Programming:Crawler:OpenWPM]], [[Programming:Crawler:Tracker Radar Collector]], [[Programming:Crawler:PageGraph]], etc.     * [[Programming:Crawler|Comparison of crawling libraries]] such as [[Programming:Crawler:OpenWPM]], [[Programming:Crawler:Tracker Radar Collector]], [[Programming:Crawler:PageGraph]], etc.
     * [[Programming:Crawler Detection|When the website notices your crawler]] — bot management, headless and CDP leaks, CAPTCHAs and challenge pages, and what a blocked crawl does to the headline number     * [[Programming:Crawler Detection|When the website notices your crawler]] — bot management, headless and CDP leaks, CAPTCHAs and challenge pages, and what a blocked crawl does to the headline number
 +    * [[Programming:Crawler:LLM Agents|Crawling with an LLM agent]] — Browser Use, BrowserGym/AgentLab and the Computer-Use APIs as a measurement instrument: what five 2026 papers measured about completion rates and per-site cost, what nobody has measured about run-to-run variance, and why this is not yet the default
 +    * [[Programming:Filter Lists|Filter lists]] — EasyList, EasyPrivacy, Disconnect and the regional lists as the field's shared labelling instrument //and// its ground truth: which list, which commit, what they miss off the anglophone desktop web, and why "blocked implies tracker" is circular
     * [[Programming:Multilingual support]]     * [[Programming:Multilingual support]]
     * [[Programming:Interaction|Interaction with websites]]     * [[Programming:Interaction|Interaction with websites]]
Line 31: Line 33:
   * [[Privacy]]   * [[Privacy]]
     * Classifying [[Privacy:Requests|Web requests]], [[Privacy:Cookies]], [[Privacy:Fingerprinting]], or [[Privacy:JavaScript]]     * Classifying [[Privacy:Requests|Web requests]], [[Privacy:Cookies]], [[Privacy:Fingerprinting]], or [[Privacy:JavaScript]]
 +    * [[Privacy:Browser storage|Storage beyond cookies]] — localStorage, IndexedDB, service workers and the caches: what your crawler records, what it silently does not, and why "we cleared cookies" is not a reset
     * [[Privacy:Server side tracking|Measuring server-side tracking]] — when the tracking request never reaches the tracker, and every request-level method quietly stops working     * [[Privacy:Server side tracking|Measuring server-side tracking]] — when the tracking request never reaches the tracker, and every request-level method quietly stops working
     * [[Privacy:Cookie syncing|Measuring cookie and ID syncing]] — how third parties learn that their two identifiers are the same person, and the crawl configuration that decides whether you see it at all     * [[Privacy:Cookie syncing|Measuring cookie and ID syncing]] — how third parties learn that their two identifiers are the same person, and the crawl configuration that decides whether you see it at all
start.1787967701.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki