User Tools

Site Tools


privacy:javascript

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
privacy:javascript [2026/08/06 22:48] – Name the unit and denominator in the opening warning box's own example. Authored by Claude. karel.kubicek.claudeprivacy:javascript [2026/08/21 14:50] (current) – Boxes: <wrap> renders a span, use uppercase <WRAP>; one box per list, not per bullet karel.kubicek.claude
Line 57: Line 57:
 ===== What This Literature Actually Is ===== ===== What This Literature Actually Is =====
  
-Everything in this section comes from a structured extraction over **4,322 full-text papers** from CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf and IEEE S&P, 2010–2024, one record per paper with a verbatim evidence quote per claim. The population here is **160 papers that analyse or classify JavaScript running in a browser** — how that population was built, and what it misses, is at [[#Reproducing These Figures]].+Everything in this section comes from a structured extraction over **5,859 full-text papers** from CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf and IEEE S&P, 2010–2026, one record per paper with a verbatim evidence quote per claim. The 2025 and 2026 venue-years are provisional — CCS and IMC 2026 have not been held and two more 2026 venue-years are incompletely selected. The population here is **206 papers that analyse or classify JavaScript running in a browser** — how that population was built, and what it misses, is at [[#Reproducing These Figures]].
  
 ==== Neither obvious search handle finds it ==== ==== Neither obvious search handle finds it ====
Line 64: Line 64:
  
 ^ Handle ^ Papers ^ Of which measure the web ^ Share ^ ^ Handle ^ Papers ^ Of which measure the web ^ Share ^
-| Used or produced a ''program-analysis'' tool | 959 215 22.4% | +| Used or produced a ''program-analysis'' tool | 1,385 299 21.6% | 
-| ''studyTypes'' includes ''code-or-binary-analysis'' | 1,063 249 | 23.4% |+| ''studyTypes'' includes ''code-or-binary-analysis'' | 1,484 347 | 23.4% |
  
-And the program-analysis tools these venues actually use are not JavaScript tools: LLVM (66 papers), Soot (61), IDA Pro (36), FlowDroid (34), angr (28), Ghidra (24), Androguard (23), Z3 (22), Apktool (21). **Esprima, at 15 papers, is the highest-ranked JavaScript parser in the whole corpus.** In a broad security corpus, "program analysis" means binaries, Android apps and smart contracts; web-script analysis is a small minority inside it, and shares almost no toolchain with the rest.+And the program-analysis tools these venues actually use are not JavaScript tools: LLVM (101 papers), Soot (81), IDA Pro (62), FlowDroid (48), angr (39), Z3 (38), Ghidra (38), CodeQL (34), Androguard (33), Apktool (30). **Esprima, at 23 papers, is the highest-ranked JavaScript parser in the whole corpus.** In a broad security corpus, "program analysis" means binaries, Android apps and smart contracts; web-script analysis is a small minority inside it, and shares almost no toolchain with the rest.
  
-==== Where the 160 papers are ====+==== Where the 206 papers are ====
  
 ^ Venue ^ Corpus papers ^ JS-analysis papers ^ Share of venue ^ ^ Venue ^ Corpus papers ^ JS-analysis papers ^ Share of venue ^
-| USENIX Security | 1,117 39 | 3.5% | +| USENIX Security | 1,410 44 | 3.1% | 
-| TheWebConf | 713 34 | 4.8% | +| TheWebConf | 843 41 | 4.9% | 
-| CCS | 889 29 | 3.3% | +| CCS | 990 33 | 3.3% | 
-| IMC | 559 20 | 3.6% | +| IMC | 638 24 | 3.8% | 
-PETS 355 18 5.1% | +IEEE S&767 22 2.9% | 
-| NDSS | 419 15 | 3.6% | +| NDSS | 701 21 | 3.0% | 
-IEEE S&270 | 1.9% |+PETS 510 21 4.1% |
  
 ^ Period ^ Corpus papers ^ JS-analysis papers ^ Per 1,000 corpus papers ^ ^ Period ^ Corpus papers ^ JS-analysis papers ^ Per 1,000 corpus papers ^
-| 2010–2013 | 468 11 23.+| 2010–2013 | 511 14 27.
-| 2014–2017 | 711 32 45.+| 2014–2017 | 769 34 44.
-| 2018–2021 | 1,403 58 41.+| 2018–2021 | 1,439 61 42.
-| 2022–2024 | 1,740 59 33.|+| 2022–2024 | 1,955 68 34.8 | 
 +| 2025–2026 //(provisional)// | 1,185 | 29 | 24.5 |
  
-The field grew sharply into the mid-2010s and has been flat-to-declining as a share of these venues since. That is not a decline in importance — it is the topic being absorbed into tracking, fingerprinting and supply-chain papers that no longer describe themselves as JavaScript analysis.+The field grew sharply into the mid-2010s and has been declining as a share of these venues since. That is not a decline in importance — it is the topic being absorbed into tracking, fingerprinting and supply-chain papers that no longer describe themselves as JavaScript analysis.
  
 ==== What they are about ==== ==== What they are about ====
Line 92: Line 93:
 Ranked, not measured: ''detection.phenomenon'' is one of the least reproducible fields in the extraction (~20% exact-string agreement between independent runs), so this is a ranking with a printed residue, never a set of percentages. Ranked, not measured: ''detection.phenomenon'' is one of the least reproducible fields in the extraction (~20% exact-string agreement between independent runs), so this is a ranking with a printed residue, never a set of percentages.
  
-^ Research family ^ Papers ^ Share of 160 +^ Research family ^ Papers ^ Share of 206 
-| Tracking- and advertising-script classification | 25 15.6% | +| Tracking- and advertising-script classification | 29 14.1% | 
-| Malicious-script and cloaking detection | 21 13.1% | +| Fingerprinting-script detection | 24 | 11.7% | 
-| Third-party libraries, inclusion and supply chain((The oldest strand of this family and still the most-cited entry point: Lauinger et al. {[lauinger2017_thou]} found **87.7% of Alexa top-75K sites and 46.5% of .com sites** using at least one of 72 catalogued JavaScript libraries, and **37.8% of Alexa sites running at least one version with a known vulnerability** (9.7% two or more).)) | 20 | 12.5% | +| Malicious-script and cloaking detection | 24 11.7% | 
-| Fingerprinting-script detection | 19 | 11.9% | +| Third-party libraries, inclusion and supply chain((The oldest strand of this family and still the most-cited entry point: Lauinger et al. {[lauinger2017_thou]} found **87.7% of Alexa top-75K sites and 46.5% of .com sites** using at least one of 72 catalogued JavaScript libraries, and **37.8% of Alexa sites running at least one version with a known vulnerability** (9.7% two or more).)) | 24 | 11.7% | 
-| Web-API usage measurement | 16 | 10.0% | +| Web-API usage measurement | 22 | 10.7% | 
-| Client-side vulnerabilities (XSS, CSP, taint flows) | 12 7.5% | +| Client-side vulnerabilities (XSS, CSP, taint flows) | 18 8.7% | 
-| Browser-extension scripts | 4.4% | +| Browser-extension scripts | 3.9% | 
-| Data leakage by scripts | | 3.8% | +| Data leakage by scripts | | 3.9% | 
-| Script change and identity over time | | 3.8% | +| Script performance, size and dead code | | 3.9% | 
-| Cryptojacking | 5 | 3.1% | +| Script change and identity over time | | 3.4% | 
-| WebAssembly (non-mining) | | 2.5% | +| Cryptojacking | 5 | 2.4% | 
-| Script performance, size and dead code | 3 | 1.9% |+| WebAssembly (non-mining) | | 2.4% |
  
-45 detection tuples did not fold into any family and are printed by the report script rather than dropped.+65 detection tuples did not fold into any family and are printed by the report script rather than dropped — a slightly larger share of a larger population than the 45 of the earlier corpus, so this fold has aged well where others have not.
  
-The shape to notice: **the privacy reader's own family is the largest but is barely a quarter of the field.** If you search these venues for "JavaScript detection" you will mostly get client-side vulnerability and malware papers, which use the same parsers and the same instrumented browsers on a different question. They are worth reading for method and misleading as related work.+The shape to notice: **the privacy reader's own two families — tracking/advertising classification and fingerprinting-script detection — lead the ranking, and together they are barely a quarter of the field (53 of 206).** If you search these venues for "JavaScript detection" you will mostly get client-side vulnerability and malware papers, which use the same parsers and the same instrumented browsers on a different question. They are worth reading for method and misleading as related work.
  
 ==== Anchor papers to read first ==== ==== Anchor papers to read first ====
Line 119: Line 120:
  
 <WRAP important> <WRAP important>
-This corpus ends in 2024. A ranking of what the 2010–2024 literature //did// is a fact about the literature, not advice about what to do now. The table below dates each method and states its status as of **2026-08-06**; the rows marked //current// were checked against post-2024 work outside the corpus.+This corpus now reaches 2026, but its 2025 and 2026 venue-years are provisional and thin. A ranking of what the 2010–2026 literature //did// is a fact about the literature, not advice about what to do now. The table below dates each method and states its status as of **2026-08-06**; the rows marked //current// were checked against post-2024 work outside the corpus, and the tool table further down now gives partial corpus support for two of them.
 </WRAP> </WRAP>
  
Line 148: Line 149:
 **LLM-based classification of web scripts is, as of 2026-08-06, essentially absent from the peer-reviewed literature.** A targeted search across PETS 2025/2026, USENIX Security 2025, NDSS 2025/2026, IMC 2025, TheWebConf 2025/2026, CCS 2025 and arXiv found no paper that classifies web scripts as trackers with a language model, or that uses one to summarise a script's privacy-relevant behaviour. The nearest work is adjacent rather than on-point: LLM-aided **deobfuscation** feeding a graph classifier for JavaScript //malware//,((//Breaking Obfuscation: Cluster-Aware Graph with LLM-Aided Recovery for Malicious JavaScript Detection//, [[https://arxiv.org/abs/2507.22447|arXiv:2507.22447]], 2025.)) LLM screening of malicious npm packages, and ''humanify'', which uses a model only to //suggest identifier names// during de-minification.(([[https://github.com/jehna/humanify|github.com/jehna/humanify]], v3.1.1, checked 2026-08-06. The AST rewrite is done by ''oxc''; the model only proposes names.)) **LLM-based classification of web scripts is, as of 2026-08-06, essentially absent from the peer-reviewed literature.** A targeted search across PETS 2025/2026, USENIX Security 2025, NDSS 2025/2026, IMC 2025, TheWebConf 2025/2026, CCS 2025 and arXiv found no paper that classifies web scripts as trackers with a language model, or that uses one to summarise a script's privacy-relevant behaviour. The nearest work is adjacent rather than on-point: LLM-aided **deobfuscation** feeding a graph classifier for JavaScript //malware//,((//Breaking Obfuscation: Cluster-Aware Graph with LLM-Aided Recovery for Malicious JavaScript Detection//, [[https://arxiv.org/abs/2507.22447|arXiv:2507.22447]], 2025.)) LLM screening of malicious npm packages, and ''humanify'', which uses a model only to //suggest identifier names// during de-minification.(([[https://github.com/jehna/humanify|github.com/jehna/humanify]], v3.1.1, checked 2026-08-06. The AST rewrite is done by ''oxc''; the model only proposes names.))
  
-<wrap todo>Treat this as an opportunity, not a settled answer. If you are planning an LLM-based script classifier, you are not late — but you also have no baseline to cite, so budget for building one, and for the reviewer question about cost, reproducibility and prompt/version drift that this page cannot yet answer for you.</wrap>+<WRAP todo>Treat this as an opportunity, not a settled answer. If you are planning an LLM-based script classifier, you are not late — but you also have no baseline to cite, so budget for building one, and for the reviewer question about cost, reproducibility and prompt/version drift that this page cannot yet answer for you.</WRAP>
  
 ==== Two 2025 results that change how you design a crawl ==== ==== Two 2025 results that change how you design a crawl ====
Line 157: Line 158:
 ===== Ground Truth, Which Is This Field's Weakest Link ===== ===== Ground Truth, Which Is This Field's Weakest Link =====
  
-Of the 160 papers, **153 record at least one classification task**. What they classify //with//:+Of the 206 papers, **198 record at least one classification task**. What they classify //with//:
  
-^ ''classification.method'' ^ Papers ^ Share of 153 +^ ''classification.method'' ^ Papers ^ Share of 198 
-| heuristic-rules | 80 52.3% | +| heuristic-rules | 108 54.5% | 
-| manual-labelling | 54 | 35.3% | +| manual-labelling | 70 | 35.4% | 
-| third-party-service | 47 30.7% | +| third-party-service | 53 26.8% | 
-| blocklist | 38 24.8% | +| blocklist | 52 26.3% | 
-| curated-database | 31 | 20.3% | +| curated-database | 40 | 20.2% | 
-| supervised-ml | 26 17.0% | +| dynamic-analysis | 33 | 16.7% | 
-| regex-or-signature | 24 15.7% | +| supervised-ml | 31 15.7% | 
-dynamic-analysis | 21 13.7% | +| regex-or-signature | 27 13.6% | 
-static-analysis | 16 10.5% | +static-analysis | 25 12.6% | 
-unsupervised-ml | 11 | 7.2% | +graph-analysis | 13 6.6% | 
-graph-analysis | 5.9% | +other | 11 | 5.6% | 
-other 5.9% |+unsupervised-ml 11 | 5.6% | 
 +**llm** **2** **1.0%** |
  
-And what they treat as truth. These 153 papers produce **278 distinct free-text ground-truth strings**, folded here into families; 52 tuples did not fold and are printed by the report script.+The ''llm'' row is new: on the 4,322-paper corpus this enum never fired for a JavaScript-classification task at all. Two papers is not a trend, and it does not contradict the finding below that no peer-reviewed paper yet classifies web scripts //as trackers// with a language model.
  
-^ Ground-truth family ^ Papers ^ Share of 153 +And what they treat as truth. These 198 papers produce **351 distinct free-text ground-truth strings**, folded here into families; 80 tuples did not fold and are printed by the report script. 
-| **Authors' own manual inspection** | 84 | **54.9%** | + 
-| Filter list or tracker database | 36 23.5% | +^ Ground-truth family ^ Papers ^ Share of 198 
-| Malware / phishing blacklist service | 11 7.2% | +| **Authors' own manual inspection** | 102 | **51.5%** | 
-| Vulnerability database | | 4.6% | +| Filter list or tracker database | 44 22.2% | 
-Synthetic or seeded ground truth | 5 | 3.3% | +| Malware / phishing blacklist service | 13 6.6% | 
-| Prior published dataset or labels | 4 | 2.6% | +| Synthetic or seeded ground truth | 9 | 4.5% | 
-| Spec or documentation | 3 | 2.0% | +| Vulnerability database | | 4.5% | 
-| Library signature catalogue | 1 | 0.7% | +Spec or documentation | 5 | 2.5% | 
-| Recruited or external annotators | 1 | 0.7% |+| Prior published dataset or labels | 4 | 2.0% | 
 +| Library signature catalogue | 1 | 0.5% | 
 +| Recruited or external annotators | 1 | 0.5% |
  
 <WRAP important> <WRAP important>
-**The most common ground truth for "is this script a tracker" is the authors reading the script.** That is a defensible choice at small scale and it is what most of this field does — but it means the labels are unpublished, unaudited and, in most cases, produced by the same people who built the classifier. Only 36 of 153 papers anchor to a filter list or tracker database, and exactly one used annotators from outside the author team.+**The most common ground truth for "is this script a tracker" is the authors reading the script.** That is a defensible choice at small scale and it is what most of this field does — but it means the labels are unpublished, unaudited and, in most cases, produced by the same people who built the classifier. Only 44 of 198 papers anchor to a filter list or tracker database, and exactly one used annotators from outside the author team.
 </WRAP> </WRAP>
  
-Validation, counted per paper rather than per tuple: **115 of 153 (75.2%) report some validation** for at least one classification, and **38 (24.8%) report none anywhere**. Where validation exists it is overwhelmingly more manual inspection (102 papers, 66.7%); comparison to another method 22 (14.4%), cross-validation 14 (9.2%), a held-out test set (5.9%).+Validation, counted per paper rather than per tuple: **146 of 198 (73.7%) report some validation** for at least one classification, and **52 (26.3%) report none anywhere**. Where validation exists it is overwhelmingly more manual inspection (129 papers, 65.2%); comparison to another method 33 (16.7%), cross-validation 16 (8.1%), a held-out test set 12 (6.1%).
  
 ==== If you use a filter list as ground truth ==== ==== If you use a filter list as ground truth ====
Line 222: Line 226:
 | Static AST / bytecode | what the code //could// do, on code you never executed | what it actually did, and anything behind ''eval'' | | Static AST / bytecode | what the code //could// do, on code you never executed | what it actually did, and anything behind ''eval'' |
  
-Which of these the 160 papers actually name:+Which of these the 206 papers actually name: 
 + 
 +^ Tool family ^ Papers ^ Share of 206 ^ 
 +| Esprima | 21 | 10.2% | 
 +| OpenWPM | 19 | 9.2% | 
 +| PageGraph | 8 | 3.9% | 
 +| Project Foxhound (taint tracking) | 8 | 3.9% | 
 +| VisibleV8 | 8 | 3.9% | 
 +| js-beautify | 7 | 3.4% | 
 +| V8 (as an analysis substrate) | 7 | 3.4% | 
 +| Closure Compiler | 5 | 2.4% | 
 +| Babel, FP-Inspector, Jalangi, WABT | 4 each | 1.9% | 
 +| JStap | 3 | 1.5% | 
 +| jsdom | 2 | 1.0% | 
 +| Acorn, AdGraph, Emscripten, JSgraph, JSNice, Khaleesi, Rhino, SpiderMonkey, TAJS, UglifyJS, unnamed deobfuscator | 1 each | 0.5% |
  
-^ Tool family ^ Papers ^ Share of 160 ^ +Two rows moved enough to matter**Esprima has overtaken OpenWPM** as the most-named tooland **Project Foxhound went from papers to 8 and PageGraph from 4 to 8** — the taint-tracking and page-graph instruments the //Methods// table above calls current are the ones the 2025–2026 papers actually picked upThat is the rare case where the corpus confirms a currency judgement instead of only dating it.
-| OpenWPM | 17 | 10.6% | +
-Esprima | 16 | 10.0% | +
-| VisibleV8 | 8 | 5.0% | +
-| js-beautify | 7 | 4.4% | +
-| V8 (as an analysis substrate) | 5 | 3.1% | +
-| Closure Compiler | 4 | 2.5% | +
-| FP-Inspector | 4 | 2.5% | +
-| PageGraph | 4 | 2.5% | +
-| WABT | 3 | 1.9% | +
-| Babel, jsdom, JStap, Project Foxhound each | 1.3% | +
-| Acorn, AdGraph, Emscripten, Jalangi, JSgraph, JSNice, Khaleesi, Rhino, SpiderMonkey, TAJS, UglifyJS, unnamed deobfuscator | 1 each | 0.6% |+
  
-**91 of 160 papers (56.9%) name no JavaScript-analysis tool at all** — they wrote their own parser, regexes or instrumentation and did not name it. That is the field's reproducibility problem in one number.+**120 of 206 papers (58.3%) name no JavaScript-analysis tool at all** — they wrote their own parser, regexes or instrumentation and did not name it. That is the field's reproducibility problem in one number.
  
 ==== Maintenance status, because half of these are frozen ==== ==== Maintenance status, because half of these are frozen ====
Line 270: Line 277:
 ===== Crawl Methodology, Against the Corpus ===== ===== Crawl Methodology, Against the Corpus =====
  
-Of the 160 papers, 128 (80.0%) ran a crawl and 126 recorded a configuration. They report better than the corpus average on every axis — and still leave the two axes that matter most for this topic largely unstated.+Of the 206 papers, 165 (80.1%) ran a crawl and 163 recorded a configuration. They report better than the corpus average on every axis — and still leave the two axes that matter most for this topic largely unstated.
  
-^ ''crawlConfig'' field ^ States a value ^ Share of 126 ^ All 829 crawling papers ^ +^ ''crawlConfig'' field ^ States a value ^ Share of 163 ^ All 1,080 papers with a crawl configuration 
-| Statefulness | 36 28.6% | 19.9% | +| Statefulness | 45 27.6% | 20.3% | 
-| Interaction depth | 112 88.9% | 78.6% | +| Interaction depth | 147 90.2% | 77.9% | 
-| Consent action | 65 51.6% | 32.6% | +| Consent action | 79 48.5% | 32.3% | 
-| Headless or headful | 28 22.2% | 13.4% | +| Headless or headful | 31 19.0% | 13.0% | 
-| Authentication | 104 | 82.5% | 71.5% | +| Authentication | 134 | 82.2% | 72.1% | 
-| Browser named | 101 80.2% | 47.6% |+| Browser named | 132 81.0% | 49.0% |
  
 <WRAP important> <WRAP important>
-**22.2% state headless-or-headful, and for this topic that is not a formality.** Two measurements in the corpus quantify the cost:+**19.0% state headless-or-headful, and for this topic that is not a formality.** Two measurements in the corpus quantify the cost:
  
   * Jueckstock et al. {[jueckstock2021_realistic]} found about **10% of script families show consistent browser-configuration bias**, and traced one concretely: Crazyegg's script invoked **fewer than 15 browser APIs under a naive crawl but nearly 60 under a stealth crawl**, because a function named ''uaBot'' short-circuits on bot detection. A naive crawl does not under-count that script — it observes a **different program**.   * Jueckstock et al. {[jueckstock2021_realistic]} found about **10% of script families show consistent browser-configuration bias**, and traced one concretely: Crazyegg's script invoked **fewer than 15 browser APIs under a naive crawl but nearly 60 under a stealth crawl**, because a function named ''uaBot'' short-circuits on bot detection. A naive crawl does not under-count that script — it observes a **different program**.
Line 291: Line 298:
 **The consent row deserves its own sentence.** On EU-facing sites a large part of the advertising and analytics stack is loaded by the consent management platform //after// a consent click, so a crawl that never interacts with the banner measures a different script population from one that accepts — and a third one from one that rejects. Half these papers do not say which they did. Decide deliberately, state it, and see [[Privacy:Consent]] for how to drive the interaction; the same choice is what makes [[Privacy:Cookies]] figures comparable or not. **The consent row deserves its own sentence.** On EU-facing sites a large part of the advertising and analytics stack is loaded by the consent management platform //after// a consent click, so a crawl that never interacts with the banner measures a different script population from one that accepts — and a third one from one that rejects. Half these papers do not say which they did. Decide deliberately, state it, and see [[Privacy:Consent]] for how to drive the interaction; the same choice is what makes [[Privacy:Cookies]] figures comparable or not.
  
-Artifact release is a bright spot: **46 of the 59 papers from 2022–2024 (78.0%) released an artifact link, against 64.7% for the corpus over the same years.** Only 13 of 160 (8.1%) assess a law — GDPR in 11 — against 6.1% corpus-wide, and 19 (11.9%) recruited participants.+Artifact release is a bright spot: **53 of the 68 papers from 2022–2024 (77.9%) released an artifact link, against 65.0% for the corpus over the same years**, and 27 of 29 (93.1%) in the provisional 2025–2026 window against 76.5%. Only 19 of 206 (9.2%) assess a law — GDPR in 17 — against 6.9% corpus-wide (402 of 5,859), and 24 (11.7%) recruited participants.
  
 ===== What to Report ===== ===== What to Report =====
Line 303: Line 310:
   - **The crawl.** Browser and version, headless or headful, stateful or stateless, consent action, interaction depth, vantage point, and — because of the results above — whether you checked for divergent behaviour under a stealth or headful configuration.   - **The crawl.** Browser and version, headless or headful, stateful or stateless, consent action, interaction depth, vantage point, and — because of the results above — whether you checked for divergent behaviour under a stealth or headful configuration.
   - **Handling of what you could not analyse.** Inline scripts, ''eval''-generated code, workers, WebAssembly, ''blob:''/''data:'' sources, and scripts that deleted themselves. Report the count you dropped rather than letting it vanish into the denominator.   - **Handling of what you could not analyse.** Inline scripts, ''eval''-generated code, workers, WebAssembly, ''blob:''/''data:'' sources, and scripts that deleted themselves. Report the count you dropped rather than letting it vanish into the denominator.
-  - **Reproducibility.** The script corpus if licensing allows, the feature extractor, and the trained model. This subfield is good at this — 78% since 2022 — so a paper without it stands out.+  - **Reproducibility.** The script corpus if licensing allows, the feature extractor, and the trained model. This subfield is good at this — 78% for 2022–2024 and higher since — so a paper without it stands out.
  
 ===== Reproducing These Figures ===== ===== Reproducing These Figures =====
Line 315: Line 322:
 Three families are in the tool table but deliberately excluded from the //membership// rule, because their non-analysis use is large and it was measured rather than assumed: OpenWPM alone dragged in an IPv6-scanning study, a QUIC website-fingerprinting paper and an HSTS study that used it as a plain crawler; Emscripten dragged in two papers that //compiled to// WebAssembly; SpiderMonkey dragged in RIDL, where the engine is the victim of a CPU attack. Three further papers — Spectre, Fallout and RIDL — were removed by hand after reading, with the reason recorded inline in the script: JavaScript is their exploit vector, not their object of study. Three families are in the tool table but deliberately excluded from the //membership// rule, because their non-analysis use is large and it was measured rather than assumed: OpenWPM alone dragged in an IPv6-scanning study, a QUIC website-fingerprinting paper and an HSTS study that used it as a plain crawler; Emscripten dragged in two papers that //compiled to// WebAssembly; SpiderMonkey dragged in RIDL, where the engine is the victim of a CPU attack. Three further papers — Spectre, Fallout and RIDL — were removed by hand after reading, with the reason recorded inline in the script: JavaScript is their exploit vector, not their object of study.
  
-^ Signal combination ^ Papers ^ Share of 160 +^ Signal combination ^ Papers ^ Share of 206 
-| classification only | 43 26.9% | +| classification only | 57 27.7% | 
-| detection only | 42 26.3% | +| detection only | 53 25.7% | 
-| detection + classification | 23 | 14.4% | +| detection + classification | 30 | 14.6% | 
-| tool only | 21 13.1% | +| tool only | 26 12.6% | 
-| tool + detection + classification | 12 7.5% | +| tool + detection + classification | 19 9.2% | 
-| tool + detection | 11 | 6.9% | +| tool + detection | 13 | 6.3% | 
-| tool + classification | 8 | 5.0% |+| tool + classification | 8 | 3.9% |
  
-The whole rule, including every fold and the reason for each hand exclusion, is below. It needs only ''extractions.jsonl'' and the shared ''lib.mjs'' helper; the companion ''report_javascript.mjs'' prints every figure on this page with its denominator, the residue of each fold, and the full 160-paper list.+The whole rule, including every fold and the reason for each hand exclusion, is below. It needs only ''extractions.jsonl'' and the shared ''lib.mjs'' helper; the companion ''report_javascript.mjs'' prints every figure on this page with its denominator, the residue of each fold, and the full 206-paper list. Every query behind this section is on [[provenance:privacy:javascript]]; corpus-level caveats are on [[literature:corpus]].
  
 <file javascript js_fold.mjs> <file javascript js_fold.mjs>
Line 332: Line 339:
 // JavaScript a page runs". Neither of the obvious schema handles answers it: // JavaScript a page runs". Neither of the obvious schema handles answers it:
 // //
-//   * `tools[].category == "program-analysis"` fires on 959 papers (used or+//   * `tools[].category == "program-analysis"` fires on 1,385 papers (used or
 //     produced), but that category is dominated by binary, Android and //     produced), but that category is dominated by binary, Android and
-//     smart-contract analysis (LLVM 66, Soot 61, IDA Pro 36, FlowDroid 34+//     smart-contract analysis (LLVM 101, Soot 81, IDA Pro 62, FlowDroid 48
-//     angr 28). Esprima, the highest-ranked JavaScript parser, is 15+//     angr 39). Esprima, the highest-ranked JavaScript parser, is 23
-//   * `studyTypes` includes `code-or-binary-analysis` on 1,063 papers, same+//   * `studyTypes` includes `code-or-binary-analysis` on 1,484 papers, same
 //     problem, and it is the least reproducible field in the schema (57%). //     problem, and it is the least reproducible field in the schema (57%).
 // //
Line 566: Line 573:
  
 // --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
-// Ground-truth sources for a script-classification task. 278 distinct strings+// Ground-truth sources for a script-classification task. 351 distinct strings
 // across the population, so again a RANKING. Ordered; first match wins, with // across the population, so again a RANKING. Ordered; first match wins, with
 // the named external resources ahead of the generic "manual" phrasings, // the named external resources ahead of the generic "manual" phrasings,
Line 592: Line 599:
 ==== Methodology and limitations of these figures ==== ==== Methodology and limitations of these figures ====
  
-  * **Seven venues only.** CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P, 2010–2024. IEEE S&P is only 43% retrieved, and **EuroS&PACSACRAID, AsiaCCS, CHI and SOUPS are absent entirely** — for this topic ACSAC and EuroS&P are a real hole, since a good deal of web-script security work lands there. Every claim here is a claim about those seven venues. +  * **Seven venues only**, 2010–2026with 2025 and 2026 provisional. Which venueswhich yearswhat each stage of the selection funnel costs and which venue-years are empty are on [[literature:corpus]] and are not restated here. For //this// topic ACSAC and EuroS&P are a real hole, since a good deal of web-script security work lands there. Every claim here is a claim about those seven venues. 
-  * **160 is a floor, and it has a false-positive tail.** Papers whose extraction never names a script as the object of detection are missing; conversely a handful of papers in the population (an e-voting client audit, a router-attack paper, a PHP injection-sink study) analyse JavaScript incidentally. The report script prints the full list so you can judge.+  * **206 is a floor, and it has a false-positive tail.** Papers whose extraction never names a script as the object of detection are missing; conversely a handful of papers in the population (an e-voting client audit, a router-attack paper, a PHP injection-sink study) analyse JavaScript incidentally. The report script prints the full list so you can judge.
   * **Not every field can carry a percentage.** ''crawlConfig.*'', ''legal.law'' and ''platforms'' reproduce to within a few points on a repeat extraction and carry the figures here. ''classification.method'' agrees on only **58%** of papers between two runs of the same schema over the same text, so its table above is a **rough share, not a precise figure** — a repeat extraction moves those rows. ''detection.phenomenon'', ''classification.resourceName'' and ''groundTruthSource'' agree on roughly 20% of exact strings, which is what the folding is for and why the family and ground-truth tables print their residue.   * **Not every field can carry a percentage.** ''crawlConfig.*'', ''legal.law'' and ''platforms'' reproduce to within a few points on a repeat extraction and carry the figures here. ''classification.method'' agrees on only **58%** of papers between two runs of the same schema over the same text, so its table above is a **rough share, not a precise figure** — a repeat extraction moves those rows. ''detection.phenomenon'', ''classification.resourceName'' and ''groundTruthSource'' agree on roughly 20% of exact strings, which is what the folding is for and why the family and ground-truth tables print their residue.
   * **Silence is not absence.** "Does not state whether it ran headless" means the paper did not say. These are reporting figures, not practice figures.   * **Silence is not absence.** "Does not state whether it ran headless" means the paper did not say. These are reporting figures, not practice figures.
-  * **Every quoted figure was checked against the paper's own text.** The prevalence values in the extraction are model summaries, so each number reproduced on this page was re-located in ''paper.cols.txt'' after whitespace normalisation. 0.9% of quotes in the dataset cannot be located in their source at all.+  * **Every quoted figure was checked against the paper's own text.** The prevalence values in the extraction are model summaries, so each number reproduced on this page was re-located in ''paper.cols.txt'' after whitespace normalisation. The dataset's own "0.9% of quotes cannot be located" figure was measured on the earlier 4,322-paper run and has not been re-measured. 
 +  * **Every query behind this section, the report script and its unedited output** are on [[provenance:privacy:javascript]]; corpus-level caveats are on [[literature:corpus]].
  
 ===== Open Questions ===== ===== Open Questions =====
  
-  * <wrap todo>**No public, hand-labelled corpus of tracking scripts exists.** Every current method builds its own labels from filter lists plus manual inspection, which is why cross-paper comparison is impossible. A shared benchmark would do for this field what EasyList did for request blocking.</wrap> +<WRAP todo> 
-  * <wrap todo>**LLM-based script classification is unmeasured.** No peer-reviewed paper found as of 2026-08-06. The obvious study — LLM against WebGraph, AdFlush and NoT.js on a fixed script corpus, reporting cost and version drift as well as F1 — has no baseline yet.</wrap> +  * **No public, hand-labelled corpus of tracking scripts exists.** Every current method builds its own labels from filter lists plus manual inspection, which is why cross-paper comparison is impossible. A shared benchmark would do for this field what EasyList did for request blocking. 
-  * <wrap todo>**Nobody has measured how much a headless or containerised crawler under-counts //script// classification specifically.** {[jueckstock2021_realistic]} and {[annamalai2024_fpfed]} show the gap exists for API traces and fingerprinting scripts; its size for tracking-script prevalence at scale is unknown.</wrap> +  * **LLM-based script classification is unmeasured.** No peer-reviewed paper found as of 2026-08-06. The obvious study — LLM against WebGraph, AdFlush and NoT.js on a fixed script corpus, reporting cost and version drift as well as F1 — has no baseline yet. 
-  * <wrap todo>**Function-granularity blocking has no successor paper.** NoT.js {[amjad2024_notjs]} and ByteDefender {[bahrami2025_bytedefender]} both stop at detection plus surrogate generation; nobody has measured what happens when either is deployed to real users at scale, or whether trackers adapt.</wrap> +  * **Nobody has measured how much a headless or containerised crawler under-counts //script// classification specifically.** {[jueckstock2021_realistic]} and {[annamalai2024_fpfed]} show the gap exists for API traces and fingerprinting scripts; its size for tracking-script prevalence at scale is unknown. 
-  * <wrap todo>**Cross-platform divergence is a confound in every older result.** If 20.6% of scripts execute differently by platform {[zafar2025_samescript]}, every desktop-only prevalence figure in this page's tables is a measurement of the desktop path only. Re-running any of them on mobile is a well-defined study.</wrap>+  * **Function-granularity blocking has no successor paper.** NoT.js {[amjad2024_notjs]} and ByteDefender {[bahrami2025_bytedefender]} both stop at detection plus surrogate generation; nobody has measured what happens when either is deployed to real users at scale, or whether trackers adapt. 
 +  * **Cross-platform divergence is a confound in every older result.** If 20.6% of scripts execute differently by platform {[zafar2025_samescript]}, every desktop-only prevalence figure in this page's tables is a measurement of the desktop path only. Re-running any of them on mobile is a well-defined study. 
 +</WRAP>
  
 ===== Related Pages ===== ===== Related Pages =====
Line 610: Line 620:
   * [[Privacy:Requests]] — classifying at the request/URL layer, where filter lists live, and why they cap this page's ground truth.   * [[Privacy:Requests]] — classifying at the request/URL layer, where filter lists live, and why they cap this page's ground truth.
   * [[Privacy:Cookies]] — what the scripts write; the provenance argument (a cookie set by a blocked resource) is the same idea one layer down.   * [[Privacy:Cookies]] — what the scripts write; the provenance argument (a cookie set by a blocked resource) is the same idea one layer down.
-  * [[Privacy:Fingerprinting]] — 42.4% of browser-fingerprinting papers are really detecting //scripts//, so that page and this one share a method. +  * [[Privacy:Fingerprinting]] — 39.8% of browser-fingerprinting papers are really detecting //scripts//, so that page and this one share a method. 
-  * [[Programming:Crawler]] — the instrumentation this page assumes you already have, compared in detail. Its per-tool pages ([[Programming:Crawler:OpenWPM]][[Programming:Crawler:PageGraph]][[Programming:Crawler:Foxhound]]) are promised but not yet written+  * [[Programming:Crawler]] — the instrumentation this page assumes you already have, compared in detail. Its per-tool pages [[Programming:Crawler:OpenWPM]] and [[Programming:Crawler:PageGraph]] are written; [[Programming:Crawler:Foxhound]] is still promised. 
-  * [[Programming:Stateful stateless]] — only 28.6% of these papers state it, and a stateless crawl sees first-visit script behaviour only.+  * [[Programming:Stateful stateless]] — only 27.6% of these papers state it, and a stateless crawl sees first-visit script behaviour only.
   * [[Design:Website classification]] — where script classification sits in the wider taxonomy.   * [[Design:Website classification]] — where script classification sits in the wider taxonomy.
   * [[Design:Crawling location]] — the vantage-point half of the "the site served you different code" problem.   * [[Design:Crawling location]] — the vantage-point half of the "the site served you different code" problem.
privacy/javascript.1786056536.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki