User Tools

Site Tools


security

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Next revision
Previous revision
security [2026/08/27 12:54] – Create security namespace outline: 5 proposed children with corpus support counts, 3 rejected, boundaries. Authored by Claude. karel.kubicek.claudesecurity [2026/08/27 14:08] (current) – Headers child now written: 44 measured of 176 probe hits. Authored by Claude. karel.kubicek.claude
Line 4: Line 4:
  
 <WRAP important> <WRAP important>
-**A namespace page outlines the pages inside it rather than carrying its own content.** (([[contributing]], "Namespace and page structure".)) The five children below are proposed from this corpus: four have a schema population, and [[Security:Headers]] has only a full-text upper bound until a later sitting hand-maps it. They are red links until written; this page exists so a reader landing from [[start]] is not sent into an empty namespace. Three candidates were not given a page; Google Safe Browsing was folded into phishing rather than made a sixth — see [[#Rejected, and why]].+**A namespace page outlines the pages inside it rather than carrying its own content.** (([[contributing]], "Namespace and page structure".)) The five children below are proposed from this corpus. [[Security:Web vulnerabilities]][[Security:TLS certificates]], [[Security:Phishing]] and [[Security:Headers]] are written; [[Security:VirusTotal]] remains a red link until written; this page exists so a reader landing from [[start]] is not sent into an empty namespace. Three candidates were not given a page; Google Safe Browsing was folded into phishing rather than made a sixth — see [[#Rejected, and why]].
 </WRAP> </WRAP>
  
Line 11: Line 11:
 ^ Page ^ What a fresh student needs it for ^ What the corpus can carry (2026-08-27) ^ ^ Page ^ What a fresh student needs it for ^ What the corpus can carry (2026-08-27) ^
 | [[Security:TLS certificates]] | You are about to measure HTTPS, certificates, or CT logs, and need to know which instrument (active scan, CT log, Censys) answers which question. | **50** papers declare their population unit as certificates (22 of them on the web platform). **133** used a TLS-specific instrument (Censys, ZGrab, crt.sh, CT, sslyze, Qualys SSL, … — **not** OpenSSL-the-library). A looser full-text probe for TLS and certificates hits **525** papers, **213** web — an upper bound the child page has to narrow. Start with Holz et al. {[holz2011_landscape]}, Durumeric et al. {[durumeric2013_https]}, Kotzias et al. {[kotzias2018_coming]}, the Censys paper {[durumeric2015_search]}, Let's Encrypt {[aas2019_encrypt]}, and ten years of ZMap {[durumeric2024_years]}. | | [[Security:TLS certificates]] | You are about to measure HTTPS, certificates, or CT logs, and need to know which instrument (active scan, CT log, Censys) answers which question. | **50** papers declare their population unit as certificates (22 of them on the web platform). **133** used a TLS-specific instrument (Censys, ZGrab, crt.sh, CT, sslyze, Qualys SSL, … — **not** OpenSSL-the-library). A looser full-text probe for TLS and certificates hits **525** papers, **213** web — an upper bound the child page has to narrow. Start with Holz et al. {[holz2011_landscape]}, Durumeric et al. {[durumeric2013_https]}, Kotzias et al. {[kotzias2018_coming]}, the Censys paper {[durumeric2015_search]}, Let's Encrypt {[aas2019_encrypt]}, and ten years of ZMap {[durumeric2024_years]}. |
-| [[Security:Headers]] | You are about to crawl for CSP, HSTS, X-Frame-Options, SRI or Trusted Types, and need to know which of those are still worth measuring. | Full-text probe (Content-Security-Policy, CSP plus "header" or "directive", HSTS, X-Frame-Options, SRI): **317** papers, **176** web. **That 176 is an upper bound, not a population** — it still contains bibliography hits and homographs. A schema match on header-ish tokens is not usable here — see the provenance page. The child page has to hand-map before any prevalence figure. Start with Weichselbaum et al. {[weichselbaum2016_dead]}, Roth et al. {[roth2020complex]}, Steffens et al. {[steffens2021_blockparty]}. |+| [[Security:Headers]] | You are about to crawl for CSP, HSTS, X-Frame-Options, SRI or Trusted Types, and need to know which of those are still worth measuring. | Full-text probe (Content-Security-Policy, CSP plus "header" or "directive", HSTS, X-Frame-Options, SRI): **317** papers, **176** web. Hand-mapped: **44** measured, 26 homographs, 13 citation-only. The 44 is 2.7% of 1,622 web-platform papers. Start with Weichselbaum et al. {[weichselbaum2016_dead]}, Roth et al. {[roth2020complex]}, Steffens et al. {[steffens2021_blockparty]}. |
 | [[Security:VirusTotal]] | You are about to label files, URLs or domains with VirusTotal, and need to know what a "detected" bit actually is. | **277** papers the extractor marked as used, produced, or drawing a population from VirusTotal (107 of them on the web platform). That is an upper bound on true use: the first sample of eight includes a references-only hit. **262** named it as a tool they used or produced; **254** of those as a classification-service. **64** distinct raw strings (products, feeds, thresholds and combinations, not merely spellings). Of the 209 that used it as a classifier, **86** targeted malware, 42 domains, 37 mobile apps, 16 website-category. The website-category use is already on [[design:website_classification]]; this page is the maliciousness-oracle use. Start with Peng et al. {[peng2019_opening]}. | | [[Security:VirusTotal]] | You are about to label files, URLs or domains with VirusTotal, and need to know what a "detected" bit actually is. | **277** papers the extractor marked as used, produced, or drawing a population from VirusTotal (107 of them on the web platform). That is an upper bound on true use: the first sample of eight includes a references-only hit. **262** named it as a tool they used or produced; **254** of those as a classification-service. **64** distinct raw strings (products, feeds, thresholds and combinations, not merely spellings). Of the 209 that used it as a classifier, **86** targeted malware, 42 domains, 37 mobile apps, 16 website-category. The website-category use is already on [[design:website_classification]]; this page is the maliciousness-oracle use. Start with Peng et al. {[peng2019_opening]}. |
 | [[Security:Phishing]] | You are about to crawl phishing sites or evaluate a feed, and need to know what the feed does not contain. | Schema union (detection, classification, population, or slug matching "phish"): **139** papers, **92** web. PhishTank is the named source in **21**. A full-text /phish/ sweep hits **1,097** papers — that is a fact about these being security venues, not a population. Start with Zhang et al. {[zhang2021_crawlphish]} (cloaking against anti-phishing crawlers) and Peng et al. {[peng2019_opening]} (VirusTotal's phishing engines). | | [[Security:Phishing]] | You are about to crawl phishing sites or evaluate a feed, and need to know what the feed does not contain. | Schema union (detection, classification, population, or slug matching "phish"): **139** papers, **92** web. PhishTank is the named source in **21**. A full-text /phish/ sweep hits **1,097** papers — that is a fact about these being security venues, not a population. Start with Zhang et al. {[zhang2021_crawlphish]} (cloaking against anti-phishing crawlers) and Peng et al. {[peng2019_opening]} (VirusTotal's phishing engines). |
security.1787835259.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki