User Tools

Site Tools


writing:literature_review

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
writing:literature_review [2026/08/27 12:59] – Apply focused-review fixes: search table scope, Alam/Warford/Birrell/Krawlers attribution, headless membership vs totals, oaklandsok hedge. Authored by Claude. karel.kubicek.claudewriting:literature_review [2026/08/27 13:04] (current) – Apply generic-pass fixes: schema vs text silence, phenomenon-label homograph, screening hedge, snowball distinction, tip box. Authored by Claude. karel.kubicek.claude
Line 10: Line 10:
 **A keyword-derived denominator is biased along the axis being measured.** **A keyword-derived denominator is biased along the axis being measured.**
  
-  * Searching this corpus for "fingerprint" hits **280** papers whose ''detection.phenomenon'' names one. **83 (29.6%)** are browser fingerprinting; **105 (37.5%)** are website / traffic fingerprinting, an encrypted-traffic-analysis problem that shares the word and almost none of the method. See [[Privacy:Fingerprinting]]. +  * Among papers whose ''detection.phenomenon'' names a fingerprint (**280** of 5,859), **83 (29.6%)** are browser fingerprinting; **105 (37.5%)** are website / traffic fingerprinting, an encrypted-traffic-analysis problem that shares the word and almost none of the method. See [[Privacy:Fingerprinting]]. 
-  * Searching for ''pre-regist*'' hits **62** papers. **47 (75.8%)** mean a pre-registered domain, OAuth redirect URI, FIDO device or test account — not a study preregistration. See [[Statistics:Study preregistration]]. +  * A full-text search for ''pre-regist*'' hits **62** papers. **47 (75.8%)** mean a pre-registered domain, OAuth redirect URI, FIDO device or test account — not a study preregistration. See [[Statistics:Study preregistration]]. 
-  * Searching for the word "SoK" in these seven venues finds IEEE S&P's track. **30 of the 43** extracted SoKs are IEEE S&P **(69.8%)**. CCS, IMC and TheWebConf contribute **zero**, in the index and in the extraction: they do not run that track+  * Title-start "SoK" in these seven venues finds IEEE S&P's track. **30 of the 43** extracted SoKs are IEEE S&P **(69.8%)**. CCS, IMC and TheWebConf contribute **zero**, in the index and in the extraction: they are not listed as SoK tracks on oaklandsok.github.io
-  * Searching crawled papers for "headless" cannot tell you how often crawls are headless. **140 of 1,120** crawled papers **(12.5%)** state ''crawlConfig.headless''; the other **87.5%** never said, and a keyword search never sees them.+  * Searching crawled papers for "headless" cannot tell you how often crawls are headless. **140 of 1,120** crawled papers **(12.5%)** state ''crawlConfig.headless''; the other **87.5%** did not state whether //their own crawl// was headless, and a keyword search never sees them.
  
 The rest of this page is that last bullet, applied to the way the field reviews itself. The rest of this page is that last bullet, applied to the way the field reviews itself.
 +</WRAP>
 +
 +<WRAP tip>
 +Before the corpus tables: define the related-work population by what the paper **did**, name the homograph, prefer a venue-complete list plus a hand filter when the quantity is a rate, and write down what the search cannot have seen. Sections 5–6 unpack that. The SoK counts below are why a Scholar pile is not that population.
 </WRAP> </WRAP>
  
Line 30: Line 34:
 ==== You cannot measure a reporting rate by searching for the thing being reported ==== ==== You cannot measure a reporting rate by searching for the thing being reported ====
  
-Suppose you want to know how often web crawls are run headless. You search Scholar for "headless crawler". Every hit is a paper that used the word. The papers that ran a crawl and never said "headless— **980 of the 1,120 (87.5%)** by the schema — are invisible, so the search cannot produce a rate. It can only produce a list of papers that spoke.+Suppose you want to know how often web crawls are run headless. You search Scholar for "headless crawler". Every hit is a paper that used the word. The papers that ran a crawl and did not state whether //their own crawl// was headless — **980 of the 1,120 (87.5%)** by the schema — are invisible to that search, so it cannot produce a rate. It can only produce a list of papers that spoke. (Some of those 980 still mention the word in related work: see the 28/31 membership split below.)
  
 That is not a Scholar limitation. It is what a keyword is. The population has to be defined **independently of the word**. Here the population is the 1,120 papers that ran a crawl; the word is then a column, not the row filter. That is not a Scholar limitation. It is what a keyword is. The population has to be defined **independently of the word**. Here the population is the 1,120 papers that ran a crawl; the word is then a column, not the row filter.
Line 154: Line 158:
 | "SoK" in the title | a web-measurement systematisation | **4 of 43 (9.3%)** are about measuring the deployed web. **26 (60.5%)** are TEE, binaries, fuzzing, crypto, hardware, DeFi. | | "SoK" in the title | a web-measurement systematisation | **4 of 43 (9.3%)** are about measuring the deployed web. **26 (60.5%)** are TEE, binaries, fuzzing, crypto, hardware, DeFi. |
  
-A related-work search that does not name the homograph will fill the section with the wrong literature. [[Privacy:Fingerprinting]] exists because that search is about 30% precise in this corpus. This page exists because the same shape repeats for every method word the field uses.+A related-work search that does not name the homograph will fill the section with the wrong literature. [[Privacy:Fingerprinting]] exists becausein this corpus, a ''detection.phenomenon'' label containing "fingerprint" is about **29.6%** browser fingerprinting. This page exists because the same shape repeats for every method word the field uses.
  
 ===== 2. What "SoK" Names in These Venues ===== ===== 2. What "SoK" Names in These Venues =====
Line 193: Line 197:
 Website-fingerprinting defenses {[mathews2023_critical]} is in the adjacent bucket on purpose: it is a web //traffic// SoK, and putting it in "web measurement" would re-create the homograph [[Privacy:Fingerprinting]] exists to kill. Website-fingerprinting defenses {[mathews2023_critical]} is in the adjacent bucket on purpose: it is a web //traffic// SoK, and putting it in "web measurement" would re-create the homograph [[Privacy:Fingerprinting]] exists to kill.
  
-==== In the index: 167 SoKs, most correctly dropped ====+==== In the index: 167 SoKs, most screened out of this corpus ====
  
 The 43 are not "the SoKs in these venues". They are the SoKs **this corpus selected**. The 43 are not "the SoKs in these venues". They are the SoKs **this corpus selected**.
Line 204: Line 208:
 | …extracted | 43 | 4 | | …extracted | 43 | 4 |
  
-**116 of 163** screened SoKs are out of scope for a measurement corpus, and looking at the PETS drop makes the reason obvious: PETS has **43** SoKs in the index and **2** in the extraction. The 41 that screening dropped include local differential privacy, MPC, federated-learning aggregation, e-voting, self-sovereign identity, trajectory generation. They are SoKs. They are not web measurement. The label did its job.+**116 of 163** screened SoKs were labelled out of scope for a measurement corpus. Looking at the PETS drop (a spot-check, not a re-read of all 116): PETS has **43** SoKs in the index and **2** in the extraction. The 41 not extracted include local differential privacy, MPC, federated-learning aggregation, e-voting, self-sovereign identity, trajectory generation. They are SoKs. They are not web measurement. Screening is why they are absent here; this sitting did not re-read every dropped abstract.
  
 The four selected-but-not-extracted SoKs are a Docker-container attack SoK (IEEE S&P 2024) and three USENIX Security 2026 papers that have been selected but not yet retrieved. The four with no abstract are all IEEE S&P 2026, and one of them is Rieder et al.'s tracker-detection SoK {[rieder2026_sok]}. Screening runs on abstracts; a paper with no abstract is never eligible, for a reason that has nothing to do with its topic. [[Literature:Corpus]] is the page that number lives on; it is repeated here because it deletes the paper you would otherwise have put first. The four selected-but-not-extracted SoKs are a Docker-container attack SoK (IEEE S&P 2024) and three USENIX Security 2026 papers that have been selected but not yet retrieved. The four with no abstract are all IEEE S&P 2026, and one of them is Rieder et al.'s tracker-detection SoK {[rieder2026_sok]}. Screening runs on abstracts; a paper with no abstract is never eligible, for a reason that has nothing to do with its topic. [[Literature:Corpus]] is the page that number lives on; it is repeated here because it deletes the paper you would otherwise have put first.
Line 269: Line 273:
   - **Prefer a venue-complete list plus a hand filter over a keyword**, for any question that is a //rate// (how often, what share, what is missing). Warford et al. {[warford2022_framework]} and Usman and Zappala {[usman2025_framework]} are the templates; Stafeev and Pellegrino {[stafeev2024_state]} is the template for "venue-complete, then a keyword", and this page is the measurement of what that keyword dropped.   - **Prefer a venue-complete list plus a hand filter over a keyword**, for any question that is a //rate// (how often, what share, what is missing). Warford et al. {[warford2022_framework]} and Usman and Zappala {[usman2025_framework]} are the templates; Stafeev and Pellegrino {[stafeev2024_state]} is the template for "venue-complete, then a keyword", and this page is the measurement of what that keyword dropped.
   - **If you must keyword-search — you usually must — write the query, the date, the venues, and the misses you know about.** "We searched Scholar for X on DATE, restricted to VENUES, and we know this drops papers that do Y without saying X." That sentence is the difference between a search and a method.   - **If you must keyword-search — you usually must — write the query, the date, the venues, and the misses you know about.** "We searched Scholar for X on DATE, restricted to VENUES, and we know this drops papers that do Y without saying X." That sentence is the difference between a search and a method.
-  - **Snowballing from a seed SoK inherits the SoK'search.** Citation snowballing is biased toward well-cited papers. If the quantity you want is "what the field does", well-cited is the axis being measured. Birrell et al. {[birrell2024_technical]} seeded from ten named venues and then snowballed; copy the named-venue part, and do not write "inside ten venues" as if the snowball stayed there.+  - **Snowballing from a seed SoK inherits the seed'inclusion rule, then adds its own bias.** Forward snowballing (who cited the seed) is biased toward well-cited papers; backward snowballing (whom the seed cited) is biased toward what that SoK already knew. Neither is a census of "what the field does". Birrell et al. {[birrell2024_technical]} seeded from ten named venues and then snowballed; copy the named-venue part, and do not write "inside ten venues" as if the snowball stayed there.
   - **Read the SoK, then read what its keyword cannot have included.** For crawling, that is Stafeev and Pellegrino plus the crawled papers that never said Tranco or Alexa. For tracking, Vekaria et al. plus the pages on this site whose populations were built from extraction fields rather than from "tracking".   - **Read the SoK, then read what its keyword cannot have included.** For crawling, that is Stafeev and Pellegrino plus the crawled papers that never said Tranco or Alexa. For tracking, Vekaria et al. plus the pages on this site whose populations were built from extraction fields rather than from "tracking".
-  - **Do not use "SoK" as a quality filter.** In these venues it is an IEEE S&P (and now USENIX / NDSS / PETS) track label. Most SoKs are not about your topic; most papers about your topic are not SoKs. CCS, IMC and TheWebConf will not give you one.+  - **Do not use "SoK" as a quality filter.** In these venues it is an IEEE S&P (and now USENIX / NDSS / PETS) track label. Most SoKs are not about your topic; most papers about your topic are not SoKs. CCS, IMC and TheWebConf are not currently listed as SoK tracks and contribute no SoK-titled records in this index.
  
 ===== 6. What to Report ===== ===== 6. What to Report =====
writing/literature_review.txt · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki