User Tools

Site Tools


writing:literature_review

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Next revision
Previous revision
writing:literature_review [2026/08/27 12:52] – Create writing:literature_review: keyword-derived denominator trap, SoK venue artifact, worked example vs this corpus. Authored by Claude. karel.kubicek.claudewriting:literature_review [2026/08/27 13:04] (current) – Apply generic-pass fixes: schema vs text silence, phenomenon-label homograph, screening hedge, snowball distinction, tip box. Authored by Claude. karel.kubicek.claude
Line 10: Line 10:
 **A keyword-derived denominator is biased along the axis being measured.** **A keyword-derived denominator is biased along the axis being measured.**
  
-  * Searching this corpus for "fingerprint" hits **280** papers whose ''detection.phenomenon'' names one. **83 (29.6%)** are browser fingerprinting; **105 (37.5%)** are website / traffic fingerprinting, an encrypted-traffic-analysis problem that shares the word and almost none of the method. See [[Privacy:Fingerprinting]]. +  * Among papers whose ''detection.phenomenon'' names a fingerprint (**280** of 5,859), **83 (29.6%)** are browser fingerprinting; **105 (37.5%)** are website / traffic fingerprinting, an encrypted-traffic-analysis problem that shares the word and almost none of the method. See [[Privacy:Fingerprinting]]. 
-  * Searching for ''pre-regist*'' hits **62** papers. **47 (75.8%)** mean a pre-registered domain, OAuth redirect URI, FIDO device or test account — not a study preregistration. See [[Statistics:Study preregistration]]. +  * A full-text search for ''pre-regist*'' hits **62** papers. **47 (75.8%)** mean a pre-registered domain, OAuth redirect URI, FIDO device or test account — not a study preregistration. See [[Statistics:Study preregistration]]. 
-  * Searching for the word "SoK" in these seven venues finds IEEE S&P's track. **30 of the 43** extracted SoKs are IEEE S&P **(69.8%)**. CCS, IMC and TheWebConf contribute **zero**, in the index and in the extraction: they do not run that track+  * Title-start "SoK" in these seven venues finds IEEE S&P's track. **30 of the 43** extracted SoKs are IEEE S&P **(69.8%)**. CCS, IMC and TheWebConf contribute **zero**, in the index and in the extraction: they are not listed as SoK tracks on oaklandsok.github.io
-  * Searching crawled papers for "headless" cannot tell you how often crawls are headless. **140 of 1,120** crawled papers **(12.5%)** state ''crawlConfig.headless''; the other **87.5%** never said, and a keyword search never sees them.+  * Searching crawled papers for "headless" cannot tell you how often crawls are headless. **140 of 1,120** crawled papers **(12.5%)** state ''crawlConfig.headless''; the other **87.5%** did not state whether //their own crawl// was headless, and a keyword search never sees them.
  
 The rest of this page is that last bullet, applied to the way the field reviews itself. The rest of this page is that last bullet, applied to the way the field reviews itself.
 +</WRAP>
 +
 +<WRAP tip>
 +Before the corpus tables: define the related-work population by what the paper **did**, name the homograph, prefer a venue-complete list plus a hand filter when the quantity is a rate, and write down what the search cannot have seen. Sections 5–6 unpack that. The SoK counts below are why a Scholar pile is not that population.
 </WRAP> </WRAP>
  
Line 30: Line 34:
 ==== You cannot measure a reporting rate by searching for the thing being reported ==== ==== You cannot measure a reporting rate by searching for the thing being reported ====
  
-Suppose you want to know how often web crawls are run headless. You search Scholar for "headless crawler". Every hit is a paper that used the word. The papers that ran a crawl and never said "headless— **980 of the 1,120 (87.5%)** by the schema — are invisible, so the search cannot produce a rate. It can only produce a list of papers that spoke.+Suppose you want to know how often web crawls are run headless. You search Scholar for "headless crawler". Every hit is a paper that used the word. The papers that ran a crawl and did not state whether //their own crawl// was headless — **980 of the 1,120 (87.5%)** by the schema — are invisible to that search, so it cannot produce a rate. It can only produce a list of papers that spoke. (Some of those 980 still mention the word in related work: see the 28/31 membership split below.)
  
 That is not a Scholar limitation. It is what a keyword is. The population has to be defined **independently of the word**. Here the population is the 1,120 papers that ran a crawl; the word is then a column, not the row filter. That is not a Scholar limitation. It is what a keyword is. The population has to be defined **independently of the word**. Here the population is the 1,120 papers that ran a crawl; the word is then a column, not the row filter.
Line 142: Line 146:
 </code> </code>
  
-The schema's 140/1,120 **(12.5%)** and the word's 137/1,120 **(12.2%)** disagree by three papers — related-work mentions, or a sentinel in the schema next to a word in the text. Neither number is a search you could have run in Scholar, because both start from the crawl population rather than from the word.+The schema's 140/1,120 **(12.5%)** and the word's 137/1,120 **(12.2%)** differ by three in the totals. Membership is not those three papers: **28** mention the word with schema sentinel or absent field, **31** have the schema set and no word in ''.cols'' (**59** papers disagree). Neither number is a search you could have run in Scholar, because both start from the crawl population rather than from the word.
  
 ==== Homographs make the false-positive side as large as the false-negative ==== ==== Homographs make the false-positive side as large as the false-negative ====
Line 154: Line 158:
 | "SoK" in the title | a web-measurement systematisation | **4 of 43 (9.3%)** are about measuring the deployed web. **26 (60.5%)** are TEE, binaries, fuzzing, crypto, hardware, DeFi. | | "SoK" in the title | a web-measurement systematisation | **4 of 43 (9.3%)** are about measuring the deployed web. **26 (60.5%)** are TEE, binaries, fuzzing, crypto, hardware, DeFi. |
  
-A related-work search that does not name the homograph will fill the section with the wrong literature. [[Privacy:Fingerprinting]] exists because that search is about 30% precise in this corpus. This page exists because the same shape repeats for every method word the field uses.+A related-work search that does not name the homograph will fill the section with the wrong literature. [[Privacy:Fingerprinting]] exists becausein this corpus, a ''detection.phenomenon'' label containing "fingerprint" is about **29.6%** browser fingerprinting. This page exists because the same shape repeats for every method word the field uses.
  
 ===== 2. What "SoK" Names in These Venues ===== ===== 2. What "SoK" Names in These Venues =====
Line 175: Line 179:
 | TheWebConf | 0 | 843 | 0.0% | | TheWebConf | 0 | 843 | 0.0% |
  
-**69.8%** of extracted SoKs are IEEE S&P. CCS, IMC and TheWebConf have no SoK track and contribute no SoK-titled paper to the index either — **0, 0, 0** of 16,864 index records. A Scholar search for "SoK" over "web measurement" is, in these venues, a search of Oakland's track plus whatever USENIX, NDSS and PETS have started to copy.+**69.8%** of extracted SoKs are IEEE S&P. CCS, IMC and TheWebConf are not listed as SoK tracks on oaklandsok.github.io and contribute no SoK-titled paper to the index either — **0, 0, 0** of 16,864 index records. A Scholar search for "SoK" over "web measurement" is, in these venues, a search of Oakland's track plus whatever USENIX, NDSS and PETS have started to copy.
  
 The 43 are not mostly about the web. After a hand fold (every paper, deciding sentence in the report): The 43 are not mostly about the web. After a hand fold (every paper, deciding sentence in the report):
Line 187: Line 191:
  
   * Stafeev and Pellegrino {[stafeev2024_state]}, USENIX Security 2024 — crawling algorithms.   * Stafeev and Pellegrino {[stafeev2024_state]}, USENIX Security 2024 — crawling algorithms.
-  * Birrell et al. {[birrell2024_technical]}, IEEE S&P 2024 — privacy-regulation impact; a citation snowball inside ten venues.+  * Birrell et al. {[birrell2024_technical]}, IEEE S&P 2024 — privacy-regulation impact; seeded from ten named venues, then expanded through backward and forward citation searches.
   * Blessing, Hugenroth, Anderson and Beresford {[blessing2025_authentication]}, PoPETs 2025 — web authentication and recovery; a 245-paper venue-and-citation census plus an Alexa top-300 passkey audit.   * Blessing, Hugenroth, Anderson and Beresford {[blessing2025_authentication]}, PoPETs 2025 — web authentication and recovery; a 245-paper venue-and-citation census plus an Alexa top-300 passkey audit.
-  * Alam et al. {[alam2026_philter]}, USENIX Security 2026 — AI phishing-website detectors; 55 papers from DBLP.+  * Alam et al. {[alam2026_philter]}, USENIX Security 2026 — AI phishing-website detectors; **38** papers from DBLP plus **17** from snowballing, **55** in total.
  
 Website-fingerprinting defenses {[mathews2023_critical]} is in the adjacent bucket on purpose: it is a web //traffic// SoK, and putting it in "web measurement" would re-create the homograph [[Privacy:Fingerprinting]] exists to kill. Website-fingerprinting defenses {[mathews2023_critical]} is in the adjacent bucket on purpose: it is a web //traffic// SoK, and putting it in "web measurement" would re-create the homograph [[Privacy:Fingerprinting]] exists to kill.
  
-==== In the index: 167 SoKs, most correctly dropped ====+==== In the index: 167 SoKs, most screened out of this corpus ====
  
 The 43 are not "the SoKs in these venues". They are the SoKs **this corpus selected**. The 43 are not "the SoKs in these venues". They are the SoKs **this corpus selected**.
Line 204: Line 208:
 | …extracted | 43 | 4 | | …extracted | 43 | 4 |
  
-**116 of 163** screened SoKs are out of scope for a measurement corpus, and looking at the PETS drop makes the reason obvious: PETS has **43** SoKs in the index and **2** in the extraction. The 41 that screening dropped include local differential privacy, MPC, federated-learning aggregation, e-voting, self-sovereign identity, trajectory generation. They are SoKs. They are not web measurement. The label did its job.+**116 of 163** screened SoKs were labelled out of scope for a measurement corpus. Looking at the PETS drop (a spot-check, not a re-read of all 116): PETS has **43** SoKs in the index and **2** in the extraction. The 41 not extracted include local differential privacy, MPC, federated-learning aggregation, e-voting, self-sovereign identity, trajectory generation. They are SoKs. They are not web measurement. Screening is why they are absent here; this sitting did not re-read every dropped abstract.
  
 The four selected-but-not-extracted SoKs are a Docker-container attack SoK (IEEE S&P 2024) and three USENIX Security 2026 papers that have been selected but not yet retrieved. The four with no abstract are all IEEE S&P 2026, and one of them is Rieder et al.'s tracker-detection SoK {[rieder2026_sok]}. Screening runs on abstracts; a paper with no abstract is never eligible, for a reason that has nothing to do with its topic. [[Literature:Corpus]] is the page that number lives on; it is repeated here because it deletes the paper you would otherwise have put first. The four selected-but-not-extracted SoKs are a Docker-container attack SoK (IEEE S&P 2024) and three USENIX Security 2026 papers that have been selected but not yet retrieved. The four with no abstract are all IEEE S&P 2026, and one of them is Rieder et al.'s tracker-detection SoK {[rieder2026_sok]}. Screening runs on abstracts; a paper with no abstract is never eligible, for a reason that has nothing to do with its topic. [[Literature:Corpus]] is the page that number lives on; it is repeated here because it deletes the paper you would otherwise have put first.
Line 214: Line 218:
 Of the 43 extracted SoKs, **19 (44.2%)** are not a paper-census at all — they are a system, an evaluation, or a vulnerability corpus with an SoK prefix. **24 (55.8%)** claim a literature sample; **18 of those 24 (75.0%)** state a search protocol a reader can name. Of the 43 extracted SoKs, **19 (44.2%)** are not a paper-census at all — they are a system, an evaluation, or a vulnerability corpus with an SoK prefix. **24 (55.8%)** claim a literature sample; **18 of those 24 (75.0%)** state a search protocol a reader can name.
  
-Among those 18:+Across all 43 (the table is not restricted to the 18):
  
 ^ How the paper list was built ^ Papers ^ Share of 43 ^ ^ How the paper list was built ^ Papers ^ Share of 43 ^
Line 227: Line 231:
 Two of those seven are the useful extremes. Two of those seven are the useful extremes.
  
-**Warford et al.** {[warford2022_framework]} (IEEE S&P 2022) gathered every paper from CCS, CHI, CSCW, IEEE S&P, NDSS, PETS, SOUPS and USENIX Security on DBLP — **6,534** papers — and hand-filtered to 127 candidates, then 95. No keyword. The cost is reading; the gain is that a paper which never says "at-risk user" can still be in.+**Warford et al.** {[warford2022_framework]} (IEEE S&P 2022) gathered every paper from CCS, CHI, CSCW, IEEE S&P, NDSS, PETS, SOUPS and USENIX Security on DBLP — **6,534** papers — and hand-filtered to 127 candidates, plus 12 from other sources to **139**, then **95**. No keyword. The cost is reading; the gain is that a paper which never says "at-risk user" can still be in.
  
 **Stafeev and Pellegrino** {[stafeev2024_state]} started from the same seven venues this corpus uses, 2010–2022, **7,840** papers, and then keyword-cut on Tranco, Alexa, and the combination of "top" and "site", to **1,057**. They discarded **654 (61.9%)** of those hits as not employing automated crawling, and surveyed the remaining **403**. Of those 403, **130 (32.3%)** navigate beyond a single page. **Stafeev and Pellegrino** {[stafeev2024_state]} started from the same seven venues this corpus uses, 2010–2022, **7,840** papers, and then keyword-cut on Tranco, Alexa, and the combination of "top" and "site", to **1,057**. They discarded **654 (61.9%)** of those hits as not employing automated crawling, and surveyed the remaining **403**. Of those 403, **130 (32.3%)** navigate beyond a single page.
Line 239: Line 243:
 | **neither** tranco nor alexa | 312 | 45.4% | | **neither** tranco nor alexa | 312 | 45.4% |
  
-**312 of 687 crawled papers in the same years (45.4%) mention neither list.** A related-work search that starts from "Tranco or Alexa" cannot see them: custom seed lists, zone files, CrUX (which that SoK **explicitly excluded** because rankings started in 2022; **7** of our 687 still mention CrUX), store crawls that the schema still scores as crawled. Their 403 is a census of crawls that named popular list, not a census of crawls.+**312 of 687 crawled papers in the same years (45.4%) mention neither list.** A related-work search that starts from "Tranco or Alexa" cannot see them: custom seed lists, zone files, CrUX (which that SoK **explicitly excluded** because rankings started in 2022; **7** of our 687 still mention CrUX), store crawls that the schema still scores as crawled. Their **403** is a census of crawls that matched one of their keyword criteria (Tranco, Alexa, or "top" with "site"), not census of crawlsand not a census of crawls that named a popular list.
  
 Their keyword false-positive is the other half: **654 of 1,057 (61.9%)** hits were not crawls. On our already-screened 2010–2022 extraction the analogue is **954 of 1,493 (63.9%)** papers matching a generous reading of their three keywords that are not in the crawled population. Screening does not remove the homograph; it only starts you closer. Their keyword false-positive is the other half: **654 of 1,057 (61.9%)** hits were not crawls. On our already-screened 2010–2022 extraction the analogue is **954 of 1,493 (63.9%)** papers matching a generous reading of their three keywords that are not in the crawled population. Screening does not remove the homograph; it only starts you closer.
Line 269: Line 273:
   - **Prefer a venue-complete list plus a hand filter over a keyword**, for any question that is a //rate// (how often, what share, what is missing). Warford et al. {[warford2022_framework]} and Usman and Zappala {[usman2025_framework]} are the templates; Stafeev and Pellegrino {[stafeev2024_state]} is the template for "venue-complete, then a keyword", and this page is the measurement of what that keyword dropped.   - **Prefer a venue-complete list plus a hand filter over a keyword**, for any question that is a //rate// (how often, what share, what is missing). Warford et al. {[warford2022_framework]} and Usman and Zappala {[usman2025_framework]} are the templates; Stafeev and Pellegrino {[stafeev2024_state]} is the template for "venue-complete, then a keyword", and this page is the measurement of what that keyword dropped.
   - **If you must keyword-search — you usually must — write the query, the date, the venues, and the misses you know about.** "We searched Scholar for X on DATE, restricted to VENUES, and we know this drops papers that do Y without saying X." That sentence is the difference between a search and a method.   - **If you must keyword-search — you usually must — write the query, the date, the venues, and the misses you know about.** "We searched Scholar for X on DATE, restricted to VENUES, and we know this drops papers that do Y without saying X." That sentence is the difference between a search and a method.
-  - **Snowballing from a seed SoK inherits the SoK'search.** Citation snowballing is biased toward well-cited papers. If the quantity you want is "what the field does", well-cited is the axis being measured. Birrell et al. {[birrell2024_technical]} snowball inside ten named venues and say so; copy the named-venue part.+  - **Snowballing from a seed SoK inherits the seed'inclusion rule, then adds its own bias.** Forward snowballing (who cited the seed) is biased toward well-cited papers; backward snowballing (whom the seed cited) is biased toward what that SoK already knew. Neither is a census of "what the field does". Birrell et al. {[birrell2024_technical]} seeded from ten named venues and then snowballed; copy the named-venue part, and do not write "inside ten venues" as if the snowball stayed there.
   - **Read the SoK, then read what its keyword cannot have included.** For crawling, that is Stafeev and Pellegrino plus the crawled papers that never said Tranco or Alexa. For tracking, Vekaria et al. plus the pages on this site whose populations were built from extraction fields rather than from "tracking".   - **Read the SoK, then read what its keyword cannot have included.** For crawling, that is Stafeev and Pellegrino plus the crawled papers that never said Tranco or Alexa. For tracking, Vekaria et al. plus the pages on this site whose populations were built from extraction fields rather than from "tracking".
-  - **Do not use "SoK" as a quality filter.** In these venues it is an IEEE S&P (and now USENIX / NDSS / PETS) track label. Most SoKs are not about your topic; most papers about your topic are not SoKs. CCS, IMC and TheWebConf will not give you one.+  - **Do not use "SoK" as a quality filter.** In these venues it is an IEEE S&P (and now USENIX / NDSS / PETS) track label. Most SoKs are not about your topic; most papers about your topic are not SoKs. CCS, IMC and TheWebConf are not currently listed as SoK tracks and contribute no SoK-titled records in this index.
  
 ===== 6. What to Report ===== ===== 6. What to Report =====
Line 288: Line 292:
   * Vekaria et al. {[vekaria2025_soktracking]} will need a venue update when one exists. As of 2026-08-27 it is arXiv v1 only.   * Vekaria et al. {[vekaria2025_soktracking]} will need a venue update when one exists. As of 2026-08-27 it is arXiv v1 only.
   * Rieder et al. {[rieder2026_sok]} will enter this corpus if an abstract lands in OpenAlex and the screening is re-run. Until then every "tracker-detection SoK" claim on this site has to say it is missing.   * Rieder et al. {[rieder2026_sok]} will enter this corpus if an abstract lands in OpenAlex and the screening is re-run. Until then every "tracker-detection SoK" claim on this site has to say it is missing.
-  * CCS / IMC / TheWebConf still have no SoK track. If one of them adds one, the venue table on this page is the thing that goes stale, not the argument.+  * CCS / IMC / TheWebConf are still not listed as SoK tracks on oaklandsok.github.io. If one of them adds one, the venue table on this page is the thing that goes stale, not the argument.
 </WRAP> </WRAP>
  
writing/literature_review.1787835145.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki