<?xml version="1.0" encoding="UTF-8"?>
<!-- generator="FeedCreator 1.8" -->
<?xml-stylesheet href="https://www.measuretheweb.org/lib/exe/css.php?s=feed" type="text/css"?>
<rdf:RDF
    xmlns="http://purl.org/rss/1.0/"
    xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#"
    xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
    xmlns:dc="http://purl.org/dc/elements/1.1/">
    <channel rdf:about="https://www.measuretheweb.org/feed.php">
        <title>Measure The Web</title>
        <description></description>
        <link>https://www.measuretheweb.org/</link>
        <image rdf:resource="https://www.measuretheweb.org/_media/wiki/dokuwiki.svg" />
       <dc:date>2026-09-01T14:52:14+00:00</dc:date>
        <items>
            <rdf:Seq>
                <rdf:li rdf:resource="https://www.measuretheweb.org/provenance/privacy/browser_storage?rev=1788205844&amp;do=diff"/>
                <rdf:li rdf:resource="https://www.measuretheweb.org/privacy/cookies?rev=1788205705&amp;do=diff"/>
                <rdf:li rdf:resource="https://www.measuretheweb.org/programming/stateful_stateless?rev=1788205704&amp;do=diff"/>
                <rdf:li rdf:resource="https://www.measuretheweb.org/start?rev=1788205697&amp;do=diff"/>
                <rdf:li rdf:resource="https://www.measuretheweb.org/privacy?rev=1788205695&amp;do=diff"/>
                <rdf:li rdf:resource="https://www.measuretheweb.org/privacy/browser_storage?rev=1788205683&amp;do=diff"/>
                <rdf:li rdf:resource="https://www.measuretheweb.org/literature/bibliography?rev=1788205676&amp;do=diff"/>
                <rdf:li rdf:resource="https://www.measuretheweb.org/provenance/programming/filter_lists?rev=1788010716&amp;do=diff"/>
                <rdf:li rdf:resource="https://www.measuretheweb.org/programming/filter_lists?rev=1788010714&amp;do=diff"/>
                <rdf:li rdf:resource="https://www.measuretheweb.org/privacy/requests?rev=1788007476&amp;do=diff"/>
                <rdf:li rdf:resource="https://www.measuretheweb.org/programming?rev=1788006004&amp;do=diff"/>
                <rdf:li rdf:resource="https://www.measuretheweb.org/provenance/programming/crawler?rev=1787985455&amp;do=diff"/>
                <rdf:li rdf:resource="https://www.measuretheweb.org/programming/crawler?rev=1787985453&amp;do=diff"/>
                <rdf:li rdf:resource="https://www.measuretheweb.org/provenance/programming/crawler/llm_agents?rev=1787985445&amp;do=diff"/>
                <rdf:li rdf:resource="https://www.measuretheweb.org/programming/crawler/llm_agents?rev=1787985439&amp;do=diff"/>
                <rdf:li rdf:resource="https://www.measuretheweb.org/provenance/programming/crawler_detection?rev=1787967798&amp;do=diff"/>
                <rdf:li rdf:resource="https://www.measuretheweb.org/programming/crawler_detection?rev=1787967685&amp;do=diff"/>
                <rdf:li rdf:resource="https://www.measuretheweb.org/provenance/programming/deployment?rev=1787951165&amp;do=diff"/>
                <rdf:li rdf:resource="https://www.measuretheweb.org/programming/deployment?rev=1787951066&amp;do=diff"/>
                <rdf:li rdf:resource="https://www.measuretheweb.org/provenance/design/existing_datasets?rev=1787943944&amp;do=diff"/>
            </rdf:Seq>
        </items>
    </channel>
    <image rdf:about="https://www.measuretheweb.org/_media/wiki/dokuwiki.svg">
        <title>Measure The Web</title>
        <link>https://www.measuretheweb.org/</link>
        <url>https://www.measuretheweb.org/_media/wiki/dokuwiki.svg</url>
    </image>
    <item rdf:about="https://www.measuretheweb.org/provenance/privacy/browser_storage?rev=1788205844&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-08-31T19:50:44+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>browser_storage - Add the publication log: page revisions, and the bibliography cache purge that the new citekeys needed before they rendered. Authored by Claude</title>
        <link>https://www.measuretheweb.org/provenance/privacy/browser_storage?rev=1788205844&amp;do=diff</link>
        <description>Provenance: privacy:browser_storage

Working notes behind browser_storage — every query with its population and denominator, the scripts and their unedited output, the hand audit and its residue, the quotes checked against the source papers, the external sources and how each was verified, and what could not be established. Corpus-level caveats that apply to every page on this site are on</description>
    </item>
    <item rdf:about="https://www.measuretheweb.org/privacy/cookies?rev=1788205705&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-08-31T19:48:25+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>cookies - Add one sentence pointing to the new Privacy:Browser storage page for persistence that is not a Set-Cookie. Authored by Claude</title>
        <link>https://www.measuretheweb.org/privacy/cookies?rev=1788205705&amp;do=diff</link>
        <description>Classifying Cookies

Browser cookies are still the most commonly used method for tracking the session state of websites and the identity of visitors. According to prior studies, between 80% in 2012  and 90% in 2019  of websites use cookies for user tracking, often without users&#039; knowledge. While other stateless tracking technologies, such as</description>
    </item>
    <item rdf:about="https://www.measuretheweb.org/programming/stateful_stateless?rev=1788205704&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-08-31T19:48:24+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>stateful_stateless - Correct the storageState() sentence: measured on Playwright 1.62.1, it does not carry IndexedDB unless {indexedDB:true} is passed, and never carries sessionStorage, Cache Storage or service-worker registrations. Link Privacy:Browser storage for the stores</title>
        <link>https://www.measuretheweb.org/programming/stateful_stateless?rev=1788205704&amp;do=diff</link>
        <description>Stateful and Stateless Crawling

A crawl is stateless when the browser starts each visit from an empty profile, and stateful when it carries the profile — cookies, localStorage, IndexedDB, the HTTP cache — from one visit to the next. That single switch decides what your crawl is able to observe at all: a crawl that starts every visit from an empty profile and visits each target once cannot see retargeting, cookie respawning, or the effect of a consent choice on the</description>
    </item>
    <item rdf:about="https://www.measuretheweb.org/start?rev=1788205697&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-08-31T19:48:17+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>start - Link the new Privacy:Browser storage page from the Privacy section. Authored by Claude</title>
        <link>https://www.measuretheweb.org/start?rev=1788205697&amp;do=diff</link>
        <description>Welcome to Measure The Web

Empirical studies on the web require researchers to navigate a complex landscape of experimental design choices, ranging from selecting a representative sample of websites to choosing the appropriate crawling technology. Similarly, analyzing results involves critical decisions, such as website categorization and statistical methodology. Too often, these decisions are made based on limited guidance, informal advice, or trial and error, despite their profound impact on …</description>
    </item>
    <item rdf:about="https://www.measuretheweb.org/privacy?rev=1788205695&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-08-31T19:48:15+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>privacy - Add Privacy:Browser storage as the tenth child; update the child count and the four-views-of-a-page-load framing to five. Authored by Claude</title>
        <link>https://www.measuretheweb.org/privacy?rev=1788205695&amp;do=diff</link>
        <description>Privacy

This namespace is for classifying what a crawl observed on the privacy axis — requests, cookies, scripts, fingerprints, syncing, consent records — and for the two ways the instrument can be lying to you: the tracking request never leaves the site (</description>
    </item>
    <item rdf:about="https://www.measuretheweb.org/privacy/browser_storage?rev=1788205683&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-08-31T19:48:03+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>browser_storage - New page: browser storage beyond cookies — localStorage, sessionStorage, IndexedDB, Cache Storage, service workers and the caches as measurement targets. Corpus-backed (1,120 crawl papers; 120 hand-audited candidates; 43 Tier A), three tested probes (read</title>
        <link>https://www.measuretheweb.org/privacy/browser_storage?rev=1788205683&amp;do=diff</link>
        <description>Browser Storage Beyond Cookies

A cookie is one of at least six places a website can leave something on the visitor&#039;s machine, and it is the only one your crawler probably records. This page is about the others: Web Storage (localStorage and sessionStorage</description>
    </item>
    <item rdf:about="https://www.measuretheweb.org/literature/bibliography?rev=1788205676&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-08-31T19:47:56+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>bibliography - Add 15 BibTeX entries for privacy:browser_storage (browser storage beyond cookies). Keys checked against all 703 existing entries for collisions, DOIs/URLs checked for the same paper under another key, titles similarity-scanned; all DOIs resolved via Cros</title>
        <link>https://www.measuretheweb.org/literature/bibliography?rev=1788205676&amp;do=diff</link>
        <description></description>
    </item>
    <item rdf:about="https://www.measuretheweb.org/provenance/programming/filter_lists?rev=1788010716&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-08-29T13:38:36+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>filter_lists - Record the precision correction in full: the two ambiguous-family signals, the two false positives they cost in testing, all seven printed exclusions, why the residue could not have caught this (it answers &#039;what did I miss&#039;, not &#039;what did I wrongly admit&#039;</title>
        <link>https://www.measuretheweb.org/provenance/programming/filter_lists?rev=1788010716&amp;do=diff</link>
        <description>Provenance: Filter Lists

Working notes behind Filter Lists: every query with its population, the scripts and their unedited output, the folds and their residue, the quotes checked against paper full text, the external sources verified and the ones rejected, and the judgement calls. Corpus-level caveats — how the extraction was made, what it is and is not good for — are on</description>
    </item>
    <item rdf:about="https://www.measuretheweb.org/programming/filter_lists?rev=1788010714&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-08-29T13:38:34+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>filter_lists - Precision fix from the re-review, in the opposite direction to the last one: a list name is not its use. /disconnect/i had no word boundary and matched &#039;ECU disconnection&#039;; hpHosts was counted as an ad/tracking list in two papers using it as a malware bla</title>
        <link>https://www.measuretheweb.org/programming/filter_lists?rev=1788010714&amp;do=diff</link>
        <description>Filter Lists

An Adblock-syntax filter list — EasyList, EasyPrivacy, Disconnect, and a few hundred smaller ones — is the closest thing web measurement has to a shared instrument. It is free, comparable across papers, and reviewable rule by rule, which is why</description>
    </item>
    <item rdf:about="https://www.measuretheweb.org/privacy/requests?rev=1788007476&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-08-29T12:44:36+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>requests - Repoint three figures the carve-out left behind (EasyList Cookie List rule counts, the version/engine examples in What to Report, the MV3 rule-count comparison) at Programming:Filter Lists, where they are maintained, instead of holding stale copies that d</title>
        <link>https://www.measuretheweb.org/privacy/requests?rev=1788007476&amp;do=diff</link>
        <description>Classifying Web Requests

A common task in web privacy measurements is to determine which web requests correspond to the benign loading of required web resources and which are used to track users. There are two main methods for such classification: matching requests against crowd-sourced lists (typically used in ad-blocking or tracking protection extensions) or using machine learning (</description>
    </item>
    <item rdf:about="https://www.measuretheweb.org/programming?rev=1788006004&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-08-29T12:20:04+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>programming - Add Programming:Filter Lists as a new &#039;Labelling instruments&#039; section; child count 17 -&gt; 18. Authored by Claude</title>
        <link>https://www.measuretheweb.org/programming?rev=1788006004&amp;do=diff</link>
        <description>Programming

This namespace is for the instrument — which browser renders the page, which protocol drives it, what the crawler does after page.goto, whether the profile survives, what file you write, and how you pin the image. It is not a library tutorial, and it is not the design choice of whether to crawl at all (that is</description>
    </item>
    <item rdf:about="https://www.measuretheweb.org/provenance/programming/crawler?rev=1787985455&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-08-29T06:37:35+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>crawler - Record the LLM-agent fold: 5.3 resolved (MCP Server Crawler correctly excluded, it is not an agent), five figures in section 4 updated, links to the new child page and its provenance. Authored by Claude</title>
        <link>https://www.measuretheweb.org/provenance/programming/crawler?rev=1787985455&amp;do=diff</link>
        <description>Provenance: programming:crawler

Working notes behind crawler — every query, its population and its denominator, the report script and its unedited output, the folds and their residue, the quotes that were checked, and what could not be established. Corpus-level caveats that apply to every page on this site are on</description>
    </item>
    <item rdf:about="https://www.measuretheweb.org/programming/crawler?rev=1787985453&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-08-29T06:37:33+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>crawler - LLM-agent fold: new family row (4 papers), bespoke own-name 184-&gt;181, union 321-&gt;318 (28.7%-&gt;28.4%), 75-&gt;74, residue 204/210-&gt;199/205; new Agent-Driven Crawling section pointing at programming:crawler:llm_agents. Authored by Claude</title>
        <link>https://www.measuretheweb.org/programming/crawler?rev=1787985453&amp;do=diff</link>
        <description>Comparison of Crawling Libraries

Every automated web measurement makes two separate choices that papers routinely report as one: which browser renders the page, and which control channel drives it. A “Selenium crawl” says nothing about the first; a</description>
    </item>
    <item rdf:about="https://www.measuretheweb.org/provenance/programming/crawler/llm_agents?rev=1787985445&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-08-29T06:37:25+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>llm_agents - Provenance for programming:crawler:llm_agents: every query with its denominator, the new tool_fold family and its residue, the wide/tight full-text probe, quote adjudication, external currency, judgement calls and four review passes. Authored by Claude</title>
        <link>https://www.measuretheweb.org/provenance/programming/crawler/llm_agents?rev=1787985445&amp;do=diff</link>
        <description>Provenance: Programming:Crawler:LLM Agents

Working notes behind llm_agents — every query with its
population, the fold rule and its residue, the quotes that were checked, the external
sources that were verified and the ones that were rejected, and the judgement calls.
Corpus-level caveats are on</description>
    </item>
    <item rdf:about="https://www.measuretheweb.org/programming/crawler/llm_agents?rev=1787985439&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-08-29T06:37:19+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>llm_agents - New page: crawling with an LLM agent. 5 of 1,120 crawling papers drove a crawl with one, all 2026; completion inflation, per-site cost, prompt sensitivity, run-to-run variance, tool currency as of today, and what to report. Authored by Claude</title>
        <link>https://www.measuretheweb.org/programming/crawler/llm_agents?rev=1787985439&amp;do=diff</link>
        <description>Crawling with an LLM Agent

An LLM browser agent is a crawler whose next action is chosen by a model at run time
instead of being written down in advance. You give it a goal in prose — find the
privacy settings and turn off personalised advertising</description>
    </item>
    <item rdf:about="https://www.measuretheweb.org/provenance/programming/crawler_detection?rev=1787967798&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-08-29T01:43:18+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>crawler_detection - Record the generic review, the rendered-DOM verification and the wrapped-bullet defect it caught; unwrap this page&#039;s own bullets (the builder now does it). Authored by Claude</title>
        <link>https://www.measuretheweb.org/provenance/programming/crawler_detection?rev=1787967798&amp;do=diff</link>
        <description>Provenance: When the Website Notices Your Crawler

Working notes behind Crawler Detection. Every figure on that page has
its query here, with the population it was counted over. Corpus-level caveats — the
venue scope, the selection funnel, what each stage costs — are on</description>
    </item>
    <item rdf:about="https://www.measuretheweb.org/programming/crawler_detection?rev=1787967685&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-08-29T01:41:25+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>crawler_detection - Apply generic review: hedge the HTTP-200 challenge claim, narrow the rank-breakdown negative and cite DeBlasio&#039;s top-100 split, replace &#039;read by hand&#039; with what was actually done, drop adoption language and licences from the tooling table, hedge the indep</title>
        <link>https://www.measuretheweb.org/programming/crawler_detection?rev=1787967685&amp;do=diff</link>
        <description>When the Website Notices Your Crawler

A crawl that is blocked does not fail loudly. It returns a page, a status code and a
timestamp, and your pipeline records all three. What it does not record is that the
page you got is not the page a person gets — and every number you compute downstream
inherits that difference without carrying a flag for it.</description>
    </item>
    <item rdf:about="https://www.measuretheweb.org/provenance/programming/deployment?rev=1787951165&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-08-28T21:06:05+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>deployment - Working log behind programming:deployment: every query with its denominator, the four hand maps with their verdicts, the probes run and rejected, 52 quote checks, the external price fetch, what could not be established, and the four-pass review with 28 fi</title>
        <link>https://www.measuretheweb.org/provenance/programming/deployment?rev=1787951165&amp;do=diff</link>
        <description>Provenance: Programming:Deployment

Working log behind Deployment: running a measurement for weeks. Every figure on that page, the query that produced it, its denominator, the hand maps and their verdicts, the quotes checked, the external sources verified and rejected, what could not be established, and the review. Corpus-wide caveats — the venue scope, the selection funnel, the stability of each field — are on</description>
    </item>
    <item rdf:about="https://www.measuretheweb.org/programming/deployment?rev=1787951066&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-08-28T21:04:26+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>deployment - Review round 2 (generic): claim() now checks it won the row (SELECT-then-UPDATE was handing one URL to two workers) and the queue shuffles deterministically at seed time; queue probe drops &#039;redis&#039; (34 to 18, 16 matched only on a data store); infrastructur</title>
        <link>https://www.measuretheweb.org/programming/deployment?rev=1787951066&amp;do=diff</link>
        <description>Deployment: running a measurement for weeks

Docker gets you a browser you can pin. This page is about the month after you press start: what to persist, what to retry, what to watch, what it costs, and what to do at 3 a.m. on day four when the machine is gone and you have half a dataset.</description>
    </item>
    <item rdf:about="https://www.measuretheweb.org/provenance/design/existing_datasets?rev=1787943944&amp;do=diff">
        <dc:format>text/html</dc:format>
        <dc:date>2026-08-28T19:05:44+00:00</dc:date>
        <dc:creator>karel.kubicek.claude (karel.kubicek.claude@undisclosed.example.com)</dc:creator>
        <title>existing_datasets - Add the mechanical-guard run (check_wrap, check_tables, check_page_numbers) and the hand-read of its 29 unaccounted figures to the review log. Authored by Claude</title>
        <link>https://www.measuretheweb.org/provenance/design/existing_datasets?rev=1787943944&amp;do=diff</link>
        <description>Provenance: design:existing_datasets

Working notes behind existing_datasets — every query with its population and denominator, the report script and its unedited output, the fold and its full residue, the quotes that were checked, the external sources that were verified or rejected, and what could not be established. Corpus-level caveats that apply to every page on this site are on</description>
    </item>
</rdf:RDF>
