| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| privacy:cookie_syncing [2026/08/26 19:08] – Re-review fixes: 80% not 83%, four hand-added papers not three, replace the 'approximately zero' overclaim with what the page's own data supports, state the late-addition asymmetry, add the gclid dataset row and explain the 9-of-13 artifact list. Authored karel.kubicek.claude | privacy:cookie_syncing [2026/08/29 12:13] (current) – Repoint the filter-list coverage anchor to the new Programming:Filter Lists page, where that section now lives. Authored by Claude karel.kubicek.claude |
|---|
| |
| <WRAP important> | <WRAP important> |
| **Read the ''headless'' row as the weakest one in this table, and not only because 80% of the papers do not state a value.** The extraction stores **one** evidence quote for the whole ''crawlConfig'' object, so a spot-check can confirm at most one of its fields per paper; of five quotes read by hand for this page, none supported the ''headless'' value and only one supported ''statefulness''. The one ''headless'' value that could be traced is also the ambiguous kind: {[englehardt2016online]} is recorded as headless because the paper describes launching measurement instances in a "headless" container "by using the pyvirtualdisplay library" to drive Xvfb — which is a //headful// browser on a virtual framebuffer, and behaves differently from a genuinely headless one under bot detection. The extraction is faithful to the paper's own word; the paper's own word is loose. If headless-ness matters to your argument, read the papers rather than this row. | **Read the ''headless'' row as the weakest one in this table, and not only because 80.0% of the papers do not state a value.** The extraction stores **one** evidence quote for the whole ''crawlConfig'' object, so a spot-check can confirm at most one of its fields per paper; of five quotes read by hand for this page, none supported the ''headless'' value and only one supported ''statefulness''. The one ''headless'' value that could be traced is also the ambiguous kind: {[englehardt2016online]} is recorded as headless because the paper describes launching measurement instances in a "headless" container "by using the pyvirtualdisplay library" to drive Xvfb — which is a //headful// browser on a virtual framebuffer, and behaves differently from a genuinely headless one under bot detection. The extraction is faithful to the paper's own word; the paper's own word is loose. If headless-ness matters to your argument, read the papers rather than this row. |
| </WRAP> | </WRAP> |
| |
| ==== How syncing papers classify requests ==== | ==== How syncing papers classify requests ==== |
| |
| Crossing the measuring set with the extraction's ''classification'' family: **20 of the 30** carry a tuple whose target is ''web-request'' (corpus-wide, 262 of the 4,439 papers that classified anything target ''web-request''). Of those 20, **15** use a **blocklist** to decide which of the parties involved counts as a tracker: **12** name EasyList and/or EasyPrivacy and **4** name Disconnect. That is a second, independent dependency on filter lists layered on top of the identifier heuristic, and it inherits [[Privacy:Requests#What a Filter List Misses, Measured|everything those lists miss]]. Only **9 of the 20** report any validation of that classification stronger than a sentinel. | Crossing the measuring set with the extraction's ''classification'' family: **20 of the 30** carry a tuple whose target is ''web-request'' (corpus-wide, 262 of the 4,439 papers that classified anything target ''web-request''). Of those 20, **15** use a **blocklist** to decide which of the parties involved counts as a tracker: **12** name EasyList and/or EasyPrivacy and **4** name Disconnect. That is a second, independent dependency on filter lists layered on top of the identifier heuristic, and it inherits [[Programming:Filter Lists#Coverage Holes, Measured|everything those lists miss]]. Only **9 of the 20** report any validation of that classification stronger than a sentinel. |
| |
| ===== What to Report ===== | ===== What to Report ===== |
| |
| If a reviewer is to accept a syncing figure, the paper has to answer all of these. The corpus can only speak to three of them, and on those three the record is poor: 24% of the papers do not state statefulness, 48% do not state a consent action, and 80% do not state whether the browser was headless. | If a reviewer is to accept a syncing figure, the paper has to answer all of these. The corpus can only speak to three of them, and on those three the record is poor: 24.0% of the papers do not state statefulness, 48.0% do not state a consent action, and 80.0% do not state whether the browser was headless. |
| |
| - **The unit and its denominator.** Sites? Directional flows? Unordered domain pairs? Users? Requests? Chains? And of what population — see the table in [[#Pick the Unit Before You Pick the Method]]. | - **The unit and its denominator.** Sites? Directional flows? Unordered domain pairs? Users? Requests? Chains? And of what population — see the table in [[#Pick the Unit Before You Pick the Method]]. |