| Next revision | Previous revision |
| practices:notifying_websites [2026/08/13 19:47] – New page: how to reach website operators at scale, response rates from 16 notification campaigns in the corpus, disclosure timelines, reporting checklist. Authored by Claude karel.kubicek.claude | practices:notifying_websites [2026/08/21 14:50] (current) – Boxes: <wrap> renders a span, use uppercase <WRAP>; one box per list, not per bullet karel.kubicek.claude |
|---|
| ====== Notifying Websites ====== | ====== Notifying Websites ====== |
| |
| You crawled 100k sites and found that 8,000 of them leak something. Before you write the paper you have to tell those 8,000 operators, and every venue in this field now expects you to say so in the paper.((See //Disclosure timelines and what venues expect// below: IMC, USENIX Security, IEEE S&P and CCS all require it in some form as of their 2026 calls.)) This page is about the mechanics of doing that at scale: **how you get a contact for a party you have never met, what response rate to plan for, how long to wait before publishing, and what to report so a reviewer accepts the result.** | You crawled 100k sites and found that 8,000 of them leak something. Before you write the paper you have to decide what to do about those 8,000 operators, and the major venues now expect the paper to say what you decided.((See //Disclosure timelines and what venues expect// below: IMC, USENIX Security, IEEE S&P and CCS all require it in some form as of their 2026 calls; PoPETs sets out ethical principles but no disclosure requirement, and NDSS and TheWebConf are not addressed there.)) This page is about the mechanics of doing that at scale: **how you get a contact for a party you have never met, what response rate to plan for, how long to wait before publishing, and what to report so a reviewer accepts the result.** |
| |
| It is not about one-off coordinated disclosure to a named vendor. Emailing Google's security team is a solved problem with a queue and an SLA; emailing 8,000 strangers is a measurement in its own right, with a delivery funnel, a control group and a p-value. Treat it as one. | It is not about one-off coordinated disclosure to a named vendor. Emailing Google's security team is a solved problem with a queue and an SLA; emailing 8,000 strangers is a measurement in its own right, with a delivery funnel, a control group and a p-value. Treat it as one. |
| |
| <WRAP important> | <WRAP important> |
| The single most consequential design decision on this page is the **control group**. Sites get fixed for reasons that have nothing to do with you — automatic updates, unrelated maintenance, going offline. Every notification study without a control arm has attributed some of that background remediation to its own notifications, and the one randomised controlled trial in this corpus found **no effect at all** from treatments that observational studies had reported as effective {[lone2022_sav]}. Randomise, hold back an arm, and report both. | The single most consequential design decision on this page is the **control group**. Sites get fixed for reasons that have nothing to do with you — automatic updates, unrelated maintenance, going offline. Every notification study without a control arm has attributed some of that background remediation to its own notifications — and **both randomised, control-arm experiments on network operators in this corpus found no significant effect from any treatment**, including treatments that observational studies had reported as effective {[lone2022_sav,qin2024_rov]}. Randomise, hold back an arm, and report both. Note what that does and does not mean: {[maass2021_effective]} is also a randomised controlled experiment — a full factorial design with a control group — and it measured 56.6% against 9.2% on German website owners. The lesson is not that notification never works. It is that a study without a control arm cannot tell you which of those two cases it is in. |
| </WRAP> | </WRAP> |
| |
| | ran a network scan or probe | 858 | 37.6% | 23.4% | 4.2% | 10.7% | 24.0% | | | ran a network scan or probe | 858 | 37.6% | 23.4% | 4.2% | 10.7% | 24.0% | |
| | assessed compliance with a law | 376 | 42.0% | 14.6% | 9.8% | 21.5% | 12.0% | | | assessed compliance with a law | 376 | 42.0% | 14.6% | 9.8% | 21.5% | 12.0% | |
| | | **crawled ∧ assessed a law** | **123** | 35.0% | 14.6% | **14.6%** | 15.4% | 20.3% | |
| |
| Two things to take from this table. **Network measurement is ahead of web measurement**: scanning papers notify at 61.0% (''yes''+''partial'') against 43.5% for crawling papers, and are less than half as likely to leave the question unanswered. The scanning community built the norm first, largely because Internet-wide scanning provoked complaints that forced the question. **Legal-compliance papers state a position most often** (only 12.0% ''not-stated'') and are also the most likely to say outright that they did not notify — a GDPR-violation finding across thousands of sites is a case where authors have decided, and said, that individual notification is not the right instrument. | Two things to take from this table. **Network measurement is ahead of web measurement**: scanning papers notify at 61.0% (''yes''+''partial'') against 43.5% for crawling papers, and are less than half as likely to say outright that they did not (4.2% vs 9.4%), as well as meaningfully less likely to leave it unanswered (24.0% vs 31.1%). The scanning community built the norm first, largely because Internet-wide scanning provoked complaints that forced the question. **Legal-compliance papers state a position most often** (only 12.0% ''not-stated''), and the bottom row is the population closest to whoever is reading this: 123 papers that both crawled the web and assessed a law, i.e. found a compliance violation across many sites. **14.6% of them say outright that they did not notify** — more than three times the corpus-wide 4.3%. That is usually a decision rather than negligence. A GDPR-violation finding across thousands of sites is the case where authors most often decide, and say, that individual notification is the wrong instrument, and route the finding to a regulator or to publication instead. If that is your situation, //deciding not to notify is a defensible position that you have to argue in the paper//, not an omission you can leave to the reader. |
| |
| The direction of travel is unambiguous: | The direction of travel is unambiguous: |
| |
| ^ Indicator ^ Denominator ^ 2010–2013 ^ 2014–2017 ^ 2018–2021 ^ 2022–2024 ^ 2025–2026* ^ | ^ Indicator ^ Denominator ^ 2010–2013 ^ 2014–2017 ^ 2018–2021 ^ 2022–2024 ^ 2025–2026* ^ |
| | Notified (''yes''\|''partial'') | empirical ∧ ethics | 21.3% | 37.7% | 45.9% | 52.8% | 60.3% | | | Notified (''yes'' or ''partial'') | empirical ∧ ethics | 21.3% | 37.7% | 45.9% | 52.8% | 60.3% | |
| | Said nothing (''not-stated'') | empirical ∧ ethics | 61.5% | 40.2% | 26.5% | 16.9% | 13.1% | | | Said nothing (''not-stated'') | empirical ∧ ethics | 61.5% | 40.2% | 26.5% | 16.9% | 13.1% | |
| | Notified, crawlers only | crawled ∧ ethics | 14.6% | 34.7% | 41.4% | 50.6% | 54.3% | | | Notified, crawlers only | crawled ∧ ethics | 14.6% | 34.7% | 41.4% | 50.6% | 54.3% | |
| | Contacted a regulator or CERT | empirical ∧ ethics | 0.3% | 3.0% | 4.7% | 3.4% | 2.8% | | | Contacted a regulator or CERT | empirical ∧ ethics | 0.3% | 3.0% | 4.7% | 3.4% | 2.8% | |
| |
| <wrap todo>* 2025–2026 is provisional: CCS 2026 and IMC 2026 have not been held, and IEEE S&P 2026 and TheWebConf 2026 abstracts are not in OpenAlex, so those venue-years are under-represented by construction. See [[Literature:Corpus]].</wrap> | <WRAP todo>* 2025–2026 is provisional: CCS 2026 and IMC 2026 have not been held, and IEEE S&P 2026 and TheWebConf 2026 abstracts are not in OpenAlex, so those venue-years are under-represented by construction. See [[Literature:Corpus]].</WRAP> |
| |
| The one flat line is the interesting one. **Regulator and CERT contact has not grown and remains rare — 3.3% of the 4,472 say yes, and only 7.2% of the 376 papers that assessed a law.** Given that a CERT is often the only party that can reach a whole constituency, and that EU law now obliges every member state to run one for exactly this purpose (below), this is the largest gap between what is available and what is used. | The one line that is not rising is the interesting one. **Regulator and CERT contact has not grown and remains rare — 3.3% of the 4,472 say yes, and only 7.2% of the 376 papers that assessed a law.** Given that a CERT is often the only party that can reach a whole constituency, and that EU law now obliges every member state to run one for exactly this purpose (below), this is the largest gap between what is available and what is used. |
| | |
| | And when the intermediary //is// used it is used **as well as**, not instead of: of the 147 papers that contacted a regulator or CERT, **135 (91.8%) also notified the affected party directly**. Nobody in this corpus treats "we told the CERT" as discharging the obligation. Plan the intermediary as an extra arm, not as the cheap way out of finding contacts. |
| |
| ==== Saying you notified is not saying how ==== | ==== Saying you notified is not saying how ==== |
| |
| 2,870 of the 4,472 (64.2%) give some free-text detail about their disclosure. Folding that text for any mention of a channel — the fold rules and their residue are on the provenance page — **2,303 of those 2,870 (80.2%) name no channel at all**, and 2,155 (75.1%) name no outcome. Where a channel is named it is usually a single large platform (11.7% mention Google, Apple, Meta, Microsoft, Amazon or Mozilla by name); a generic email to the operator or developer is 3.9%, a CERT 1.4%, a hosting provider 1.3%, a bug-bounty programme 1.3%, WHOIS 0.6%, a data protection authority 0.6%. | 2,870 of the 4,472 (64.2%) give some free-text detail about their disclosure, and **all 2,160 that said they notified are among them**. Scoping to those 2,160 — a channel is only a fair question of a paper that says it notified somebody((Of the 710 papers that gave a detail without saying they notified, only 273 (38.5%) are participant-consent or debriefing notes; the rest are data-handling and harm-limitation notes, and a handful even name a channel. So the exclusion is about the question being ill-posed for that group, not about its text being all one kind.)) — and folding the text for any mention of a channel (the rules and their residue are on the provenance page): |
| | |
| | **1,633 of the 2,160 (75.6%) do not say through what channel.** Where a channel does appear it is usually a single large platform named outright (14.7% mention Google, Apple, Meta, Microsoft, Amazon or Mozilla); a generic email to the operator or developer is 4.6%, a CERT 1.8%, a bug-bounty programme 1.6%, a hosting provider 1.6%, WHOIS 0.7%, a data protection authority 0.6%. And 2,155 of the 2,870 (75.1%) name no outcome either. |
| |
| And **only 7 of the 2,870 report a countable response ratio in that text** — figures like "Contacted 318 of 559 sensitive organizations; received 11 responses by submission" or "Reported 110 malicious extensions to Google; 62.7% were removed". Every rate in the next section had to be read out of the papers' own results sections, not out of their ethics sections. That is the reporting gap this page exists to close: the modal paper in this corpus says //we disclosed responsibly// and stops. | Worse for anyone trying to plan a campaign: **only 7 of the 2,870 report a countable response ratio in that text** — figures like "Contacted 318 of 559 sensitive organizations; received 11 responses by submission" or "Reported 110 malicious extensions to Google; 62.7% were removed". Every rate in the next section had to be read out of the papers' own results sections, not out of their ethics sections. That is the reporting gap this page exists to close: the modal paper in this corpus says //we disclosed responsibly// and stops. |
| |
| ===== Getting a contact: what still works in 2026 ===== | ===== Getting a contact: what still works in 2026 ===== |
| |
| This is where notification studies actually fail. Both 2016 campaigns concluded that reachability, not operator willingness, was the binding constraint {[stock2016_hey,li2016_youve]}, and the 2026 interview study of hosting providers reports the same thing a decade later: "the response and remediation rates for notified web vulnerabilities still average between 20 to 30%, and no effective alternatives have yet been identified" {[stivala2026_behind]}. | This is where notification studies actually fail. Both 2016 campaigns concluded that reachability, not operator willingness, was the binding constraint {[stock2016_hey,li2016_youve]}, and the 2026 interview study of hosting providers opens by summarising the decade since as "the response and remediation rates for notified web vulnerabilities still average between 20 to 30%, and no effective alternatives have yet been identified (e.g., still no better channel than WHOIS)" {[stivala2026_behind]}. That sentence is a related-work summary carrying citations to four earlier papers, not a figure that study measured — but it is the field's own 2026 assessment of where it stands, and it is worth knowing that a 2026 paper still writes it. |
| |
| ==== The channel that carried this literature is gone ==== | ==== The channel that carried this literature is gone ==== |
| |
| <WRAP important> | <WRAP important> |
| **Domain WHOIS is no longer a contact channel.** Both 2016 campaigns bought machine-readable WHOIS for the Alexa top million and used the registrant or technical address; it was the best channel in {[stock2016_hey]} — highest fix rate (+15.4% over control) and the lowest bounce rate, "since a valid email address is typically necessary to register a domain". Two things then removed it: | **Domain WHOIS is no longer a contact channel.** Both 2016 campaigns bought machine-readable WHOIS for the Alexa top million and used the registrant or technical address; it was the best channel in {[stock2016_hey]} — "the communication channel with the highest fix ratio (+15.4%)" relative to control, and far fewer bounces than the generic aliases, "since a valid email address is typically necessary to register a domain". Two things then removed it: |
| |
| * **GDPR, from May 2018.** ICANN's Temporary Specification told registrars to redact registrant, admin and tech contact fields. Measured over 1.2 billion WHOIS records, "over 85% [of] surveyed large WHOIS providers [were] redacting EEA records at scale" and "over 60% [of] large WHOIS data providers also redact non-EEA records" {[lu2021_whowas]}. The same paper surveyed 51 security papers using WHOIS and found 69% of them needed a field that is now redacted. | * **GDPR, from May 2018.** ICANN's Temporary Specification told registrars to redact registrant, admin and tech contact fields. Measured over 1.2 billion WHOIS records, "over 85% [of] surveyed large WHOIS providers [were] redacting EEA records at scale" and "over 60% [of] large WHOIS data providers also redact non-EEA records" {[lu2021_whowas]}. The same paper surveyed 51 security papers using WHOIS and found 69% of them needed a field that is now redacted. |
| * **The protocol itself, from 28 January 2025.** ICANN: "As of 28 January 2025, the Registration Data Access Protocol (RDAP) will be the definitive source for delivering generic top-level domain name (gTLD) registration information in place of sunsetted WHOIS services."((ICANN, //ICANN Update: Launching RDAP; Sunsetting WHOIS//, 27 January 2025, [[https://www.icann.org/en/announcements/details/icann-update-launching-rdap-sunsetting-whois-27-01-2025-en|icann.org]]. Fetched 2026-08-13. The replacement Registration Data Policy took effect for contracted parties in August 2025.)) Port-43 WHOIS for gTLDs is not a thing you should be building a 2026 pipeline on. | * **The protocol itself, from 28 January 2025.** ICANN: "As of 28 January 2025, the Registration Data Access Protocol (RDAP) will be the definitive source for delivering generic top-level domain name (gTLD) registration information in place of sunsetted WHOIS services."((ICANN, //ICANN Update: Launching RDAP; Sunsetting WHOIS//, 27 January 2025, [[https://www.icann.org/en/announcements/details/icann-update-launching-rdap-sunsetting-whois-27-01-2025-en|icann.org]]. Fetched 2026-08-13. ICANN's replacement Registration Data Policy is separately announced as "Now In Effect for Contracted Parties"; the exact commencement date could not be pinned to a primary source and is therefore not stated here.)) Port-43 WHOIS for gTLDs is not a thing you should be building a 2026 pipeline on. |
| |
| RDAP (RFC 9082/9083, discovery per RFC 9224) is the successor, but it is a **protocol** replacement, not a **data** replacement: it returns the same redacted fields, plus a documented mechanism for differentiated access. If you need the non-public registration data, ICANN's Registration Data Request Service (RDRS) is the front door, and it names "cybersecurity professionals" among those with a legitimate interest.((Same ICANN announcement. RDRS covers participating registrars only; for the rest you contact the sponsoring registrar directly.)) That is a per-request human process, so it is a channel for tens of domains, not for thousands. | RDAP (RFC 9082/9083, discovery per RFC 9224) is the successor, but it is a **protocol** replacement, not a **data** replacement: it returns the same redacted fields, plus a documented mechanism for differentiated access. If you need the non-public registration data, ICANN's Registration Data Request Service (RDRS) is the front door, and it names "cybersecurity professionals" among those with a legitimate interest.((Same ICANN announcement. RDRS covers participating registrars only; for the rest you contact the sponsoring registrar directly.)) That is a per-request human process, so it is a channel for tens of domains, not for thousands. |
| |
| ^ Channel ^ How you get it ^ What it reaches ^ Status and evidence ^ | ^ Channel ^ How you get it ^ What it reaches ^ Status and evidence ^ |
| | **''security.txt''** (RFC 9116) | ''GET https://<domain>/.well-known/security.txt'' | the security team, if there is one | RFC 9116, Informational, April 2022 — current, not obsoleted.((Verified against [[https://www.rfc-editor.org/rfc/rfc9116.txt|rfc-editor.org/rfc/rfc9116.txt]] on 2026-08-13: "Category: Informational … April 2022". Three errata exist, none touching the field list.)) Adoption is the problem: 11–16% of the Alexa top 100, 8–10% of the top 1K, "3–4% for the top 10K sites, and only a percent for the top 100K" {[poteat2021_securitytxt]} | | | **''security.txt''** (RFC 9116) | ''GET https://<domain>/.well-known/security.txt'' | the security team, if there is one | RFC 9116, Informational, April 2022 — current, not obsoleted.((Verified against [[https://www.rfc-editor.org/rfc/rfc9116.txt|rfc-editor.org/rfc/rfc9116.txt]] on 2026-08-13: "Category: Informational … April 2022". Three errata exist, none touching the field list.)) Adoption is the problem: 11–16% of the Alexa top 100, 8–10% of the top 1K, "3–4% for the top 10K sites, and only a percent for the top 100K" {[poteat2021_securitytxt]}. **That measurement is from 2021 and its ranking frame no longer exists** — Alexa was discontinued in 2022 (see [[Design:Website selection]]), the paper predates RFC 9116, and **nothing in this corpus re-measures security.txt adoption since**. Assume it has risen; you have no citeable current figure | |
| | **RIR abuse contact** | RIPE ''abuse-c:'' / ARIN Abuse POC, resolved from the IP; RIPEstat's abuse-contact-finder covers all five RIRs over HTTPS | the **hosting provider or CDN**, not the site owner | The one source that scales — free, no meaningful rate limit, annually validated by RIPE and ARIN. Still the most-mentioned inbound channel among hosting providers in 2026 {[stivala2026_behind]} | | | **RIR abuse contact** | RIPE ''abuse-c:'' / ARIN Abuse POC, resolved from the IP; RIPEstat's abuse-contact-finder covers all five RIRs over HTTPS | the **hosting provider or CDN**, not the site owner | The one source that scales — free and with no meaningful rate limit. ARIN verifies its Abuse POC annually;((ARIN NRPM §3.6: "Each of the following Points of Contact are to be verified annually … Admin, Tech, NOC, Abuse", [[https://www.arin.net/participate/policy/nrpm/|arin.net]], fetched 2026-08-13. RIPE states it works to keep abuse contacts valid but no validation cadence could be found on a RIPE primary source on 2026-08-13 — do not assume annual.)) RIPE states it keeps abuse contacts valid but publishes no cadence that could be verified from a RIPE source. Still the workhorse: "WHOIS was the most frequently mentioned channel (seven HPOs)" among 24 interviewed providers in 2026, who "confirmed that their contact is available via WHOIS, for example through ARIN or RIPE databases" {[stivala2026_behind]}. Note this is the //network// registry, which GDPR redaction did not touch — not the //domain// registrant record, which it did | |
| | **PeeringDB technical contact** | PeeringDB API, per AS | the network operator | Preferred over WHOIS by {[lone2022_sav]}, which found PeeringDB better maintained | | | **PeeringDB technical contact** | PeeringDB API, per AS | the network operator | Preferred over WHOIS by {[lone2022_sav]}: "We preferred peeringDB because it has been used in previous studies and they found the database up-to-date" — i.e. it is relaying prior work's assessment, not measuring it | |
| | **The site's own imprint / privacy-policy contact** | read it off the page | the legally responsible party — often the actual decision-maker | The highest-yield channel measured. Addresses parsed from privacy-policy and contact pages achieved **87.8% delivery against 33.8% for RFC 2142 aliases** {[utz2023_comparing]}; 9 of 11 organisations reached via a privacy-policy address resolved the issue {[elyadmani2025_keys]}. {[maass2021_effective]} collected German ''Impressum'' addresses by hand, three researchers per site | | | **The site's own imprint / privacy-policy contact** | read it off the page | the legally responsible party — often the actual decision-maker | The highest-yield channel measured. Addresses parsed from privacy-policy and contact pages achieved **87.8% delivery against 33.8% for RFC 2142 aliases** {[utz2023_comparing]}; 9 of 11 organisations reached via a privacy-policy address resolved the issue {[elyadmani2025_keys]}. {[maass2021_effective]} collected German ''Impressum'' addresses by hand, three researchers per site | |
| | **RFC 2142 role aliases** (''security@'', ''abuse@'', ''webmaster@'', ''info@'') | construct from the domain | whoever reads that mailbox, if anyone | RFC 2142 still current, still unrevised since 1997. Cheap and weak: 33.8% delivery {[utz2023_comparing]}, and for half the WordPress domains in {[stock2016_hey]} //every// alias bounced | | | **RFC 2142 role aliases** (''security@'', ''abuse@'', ''webmaster@'', ''info@'') | construct from the domain | whoever reads that mailbox, if anyone | RFC 2142 still current, still unrevised since 1997. Cheap and weak: 33.8% delivery {[utz2023_comparing]}, and for half the WordPress domains in {[stock2016_hey]} //every// alias bounced | |
| | **HackerOne Disclosure Assistance** | ''hackerone.com/disclosure-assistance'' | organisations with no disclosure policy at all | Newer than the literature: HackerOne "will work with friendly hackers on a **best-effort basis**" to verify a bug, find someone at the affected organisation and relay it.((Fetched from [[https://docs.hackerone.com/en/articles/8466632-disclosure-assistance|docs.hackerone.com]] on 2026-08-13; page dated 11 June 2024.)) It updates {[stock2016_hey]}, which discarded reward programmes because they "usually only accept and forward reports for their customers" — that premise no longer holds for HackerOne. But it is a per-report human service gated on having exhausted other options, so it is **not** a bulk channel | | | **HackerOne Disclosure Assistance** | ''hackerone.com/disclosure-assistance'' | organisations with no disclosure policy at all | Newer than the literature: HackerOne "will work with friendly hackers on a **best-effort basis**" to verify a bug, find someone at the affected organisation and relay it.((Fetched from [[https://docs.hackerone.com/en/articles/8466632-disclosure-assistance|docs.hackerone.com]] on 2026-08-13; page dated 11 June 2024.)) It updates {[stock2016_hey]}, which discarded reward programmes because they "usually only accept and forward reports for their customers" — that premise no longer holds for HackerOne. But it is a per-report human service gated on having exhausted other options, so it is **not** a bulk channel | |
| |
| The following **do not scale**, and the literature says so explicitly rather than leaving you to find out: web contact forms and telephone ({[stock2016_hey]} discarded both), and scraping ''mailto:'' links off pages (unreliable, defeated by CAPTCHAs and obfuscation, and it collects addresses nobody published for this purpose). | Scraping ''mailto:'' links off pages **does not scale** in a way the literature is explicit about: unreliable, defeated by CAPTCHAs and obfuscation, and it collects addresses nobody published for this purpose. {[stock2016_hey]} discarded it, along with web contact forms and telephone, as "not a viable option in a large-scale scenario". |
| | |
| | <WRAP info> |
| | **"Does not scale" is a function of your N, and the highest rates in this literature come from channels that do not scale.** {[sasaki2022_ics]} notified 160 ICS operators **by manual phone call** — with a dedicated team, over three and a half months — and got the person actually responsible for 212 devices, 50% stating they would mitigate and a 58% confirmed reduction in exposed devices against 13% for un-notified ones. {[maass2021_effective]} read 4,594 ''Impressum'' addresses by hand, three researchers per site, and posted paper letters. Both are far above the 5–20% band that automated email campaigns report. |
| | |
| | If your finding affects hundreds of parties rather than tens of thousands, **the manual channel is the method**, and the tooling in this section is the wrong tool. The trade is a real one: {[sasaki2022_ics]} publishes its cost per notification, and {[maass2021_effective]} spent about €5,000 on postage. Decide which you are doing before you build a pipeline. |
| | </WRAP> |
| |
| ==== A tested contact-discovery cascade ==== | ==== A tested contact-discovery cascade ==== |
| |
| Three of those channels are machine-queryable. This runs them in cost order and prints per-source coverage, because a pooled "we found contacts for N% of domains" is the figure a reviewer sends back. Failures are printed rather than swallowed: a silent zero and a missing contact look identical, and the difference understates your own coverage. | Two of those channels are machine-queryable per domain, and RDAP is worth querying even though it is mostly redacted, because what it returns tells you //which kind// of redaction you are up against. (PeeringDB is machine-queryable too, but per AS — use it when your unit is a network operator rather than a website.) This runs the three in cost order and prints per-source coverage, because a pooled "we found contacts for N% of domains" is the figure a reviewer sends back. Failures are printed rather than swallowed: a silent zero and a missing contact look identical, and the difference understates your own coverage. |
| |
| <file python find_contacts.py> | <file python find_contacts.py> |
| * scraping mailto: links off the page -- highest yield, worst measured | * scraping mailto: links off the page -- highest yield, worst measured |
| delivery, and it collects addresses nobody published for this purpose. | delivery, and it collects addresses nobody published for this purpose. |
| | |
| | Why RIPEstat rather than the Abusix Abuse Contact DB for the RIR abuse contact: |
| | Abusix is a DNS TXT lookup, free and unmetered, and is what the 2016 USENIX |
| | campaign used, so it is the better choice at real scale. RIPEstat is used here |
| | because it needs nothing but the standard library over HTTPS, which keeps this |
| | script runnable -- and therefore testable -- anywhere. Swap in Abusix once you have |
| | a DNS client and thousands of domains. |
| |
| Failures are printed, never swallowed: a silent zero looks identical to a | Failures are printed, never swallowed: a silent zero looks identical to a |
| missing contact and would understate your own coverage. | missing contact and would understate your own coverage. The one exception is the |
| | IANA bootstrap fetch, which raises: it is not a per-domain zero, it makes RDAP |
| | coverage unmeasurable for the whole batch. |
| """ | """ |
| |
| def _rdap_base(tld): | def _rdap_base(tld): |
| if not _BOOTSTRAP_CACHE: | if not _BOOTSTRAP_CACHE: |
| for tlds, urls in json.loads(_get(BOOTSTRAP))["services"]: | # Fetched once per run. A failure here is NOT a per-domain coverage zero |
| | # -- it means RDAP coverage is unmeasurable for the whole batch -- so it |
| | # raises with the URL rather than degrading into a row of dashes that |
| | # would read as "no contact found" for every domain in the sample. |
| | try: |
| | services = json.loads(_get(BOOTSTRAP))["services"] |
| | except (urllib.error.URLError, socket.timeout, OSError, json.JSONDecodeError, KeyError) as e: |
| | raise RuntimeError( |
| | f"RDAP bootstrap registry unreachable ({BOOTSTRAP}): {type(e).__name__}: {e}. " |
| | "Refusing to continue: every RDAP row would print as 'not found' and " |
| | "understate coverage. Retry, or run with RDAP excluded and say so." |
| | ) from e |
| | for tlds, urls in services: |
| for t in tlds: | for t in tlds: |
| _BOOTSTRAP_CACHE[t] = urls[0].rstrip("/") | _BOOTSTRAP_CACHE[t] = urls[0].rstrip("/") |
| if contact: | if contact: |
| hit[src].add(d) | hit[src].add(d) |
| print(f"{d:24} {src:13} {(contact or '-')[:40]:40} {note}") | print(f"{d:24} {src:13} {(contact or '--')[:40]:40} {note}") |
| print(f"\ncoverage, domains with >=1 usable contact, of {len(domains)}:") | print(f"\ncoverage, domains with >=1 usable contact, of {len(domains)}:") |
| for s in SOURCES: | for s in SOURCES: |
| google.com security.txt mailto:security@google.com expires=2030-04-01T00:00:00z | google.com security.txt mailto:security@google.com expires=2030-04-01T00:00:00z |
| google.com RDAP — response carried no email — redacted, or role-only | google.com RDAP — response carried no email — redacted, or role-only |
| google.com RIR abuse — RIPEstat: TimeoutError | google.com RIR abuse network-abuse@google.com ip=172.217.208.139 rir=arin |
| ethz.ch security.txt mailto:security@ethz.ch expires=2028-01-31T07:00:00.000Z | ethz.ch security.txt mailto:security@ethz.ch expires=2028-01-31T07:00:00.000Z |
| ethz.ch RDAP — .ch has no RDAP service in the IANA bootstrap | ethz.ch RDAP — .ch has no RDAP service in the IANA bootstrap |
| bbc.co.uk security.txt mailto:security@bbc.co.uk expires=2038-01-19T03:14:07Z | bbc.co.uk security.txt mailto:security@bbc.co.uk expires=2038-01-19T03:14:07Z |
| bbc.co.uk RDAP redacted@nominet.uk role=registrant | bbc.co.uk RDAP redacted@nominet.uk role=registrant |
| bbc.co.uk RIR abuse abuse@fastly.com ip=151.101.192.81 rir=arin | bbc.co.uk RIR abuse abuse@fastly.com ip=151.101.128.81 rir=arin |
| wikipedia.org security.txt mailto:security@wikimedia.org expires=2029-03-31T09:00:00.000Z | wikipedia.org security.txt mailto:security@wikimedia.org expires=2029-03-31T09:00:00.000Z |
| wikipedia.org RDAP — response carried no email — redacted, or role-only | wikipedia.org RDAP — response carried no email — redacted, or role-only |
| security.txt 4 80.0% | security.txt 4 80.0% |
| RDAP 1 20.0% | RDAP 1 20.0% |
| RIR abuse 4 80.0% | RIR abuse 5 100.0% |
| any source 5 100.0% | any source 5 100.0% |
| </code> | </code> |
| * **ccTLDs often have no RDAP.** ''.ch'' and ''.de'' have no service in the IANA bootstrap registry. If your sample is a national top list, RDAP coverage may be zero and it will not be your bug. | * **ccTLDs often have no RDAP.** ''.ch'' and ''.de'' have no service in the IANA bootstrap registry. If your sample is a national top list, RDAP coverage may be zero and it will not be your bug. |
| * **The RIR abuse contact reaches the infrastructure, not the site.** ''cispa.de'' resolves to its hosting provider's ''abuse@she.net''; ''bbc.co.uk'' sits behind Fastly and yields ''abuse@fastly.com''. For any CDN-fronted site the abuse contact is the CDN — and CDN-fronting is now the common case, which is a structural reason this channel has decayed since 2016. Report what fraction of your abuse contacts belong to a CDN or shared host rather than to the party you meant to reach. | * **The RIR abuse contact reaches the infrastructure, not the site.** ''cispa.de'' resolves to its hosting provider's ''abuse@she.net''; ''bbc.co.uk'' sits behind Fastly and yields ''abuse@fastly.com''. For any CDN-fronted site the abuse contact is the CDN — and CDN-fronting is now the common case, which is a structural reason this channel has decayed since 2016. Report what fraction of your abuse contacts belong to a CDN or shared host rather than to the party you meant to reach. |
| * **A transient failure looks exactly like an absent contact.** The ''RIPEstat: TimeoutError'' line is a retry, not a coverage figure. Distinguish the two or your denominator is wrong. | * **A transient failure looks exactly like an absent contact.** An earlier run of exactly this script, on exactly these five domains, printed ''RIPEstat: TimeoutError'' for ''google.com'' and so reported 80% rather than 100% RIR-abuse coverage. Nothing about the sample changed. Distinguish a network failure from a missing contact and retry it, or your denominator is wrong — and note that the script **refuses to run at all** if the IANA bootstrap registry is unreachable, because degrading that into a column of dashes would understate RDAP coverage for every domain at once. |
| |
| ===== Response and remediation rates to plan for ===== | ===== Response and remediation rates to plan for ===== |
| |
| These are the campaigns in the corpus that reported a quotable outcome, each with **its own denominator** — they are not comparable to each other, because "response", "remediation" and "fix" are defined differently in each and the populations are wildly different. Plan against the range, not the mean. The hand pass classified 22 campaigns in total; the seven not shown here reported a volume sent but no response or remediation figure, and are listed on [[provenance:practices:notifying_websites]]. | {[sasaki2022_ics]} collects the comparison itself, and it is the shortest statement of the range: "its remediation rate was approximately 18% … Our remediation rate is higher than most previous notification experiments: it was approximately 40% for cross-site scripting and a WordPress vulnerability, 33%–42% for different WordPress vulnerability, and less than 20% for DNS zone poisoning. The only campaigns that reported similar remediation rates were on publicly accessible Git repositories (78%–81%) and Heartbleed (approximately 40%–90%)." |
| | |
| | These are the campaigns in the corpus that reported a quotable outcome, each with **its own denominator** — they are not comparable to each other, because "response", "remediation" and "fix" are defined differently in each and the populations are wildly different. Plan against the range, not the mean. The hand pass classified 22 campaigns in total; 16 are below, and the six not shown reported a volume sent but no response or remediation figure. One row below — {[munteanu2025_catch22]} — is here for its channel rather than a rate, because delegating a campaign to an established notification operator is the pattern, not the number. They are listed on [[provenance:practices:notifying_websites]], along with the one row below — {[stock2018_didnt]} — whose full text is missing from the corpus and was read from the publisher's PDF instead. |
| |
| ^ Study ^ Notified ^ Channel ^ Outcome, in the paper's own terms ^ Control ^ | ^ Study ^ Notified ^ Channel ^ Outcome, in the paper's own terms ^ Control ^ |
| | {[canali2013_hosting]} TheWebConf 2013 | global and regional hosting providers | abuse notification to the provider | "50% of both the global and regional web hosting providers never replied to any of the real abuse notifications we sent" | — | | | {[canali2013_hosting]} TheWebConf 2013 | global and regional hosting providers | abuse notification to the provider | "50% of both the global and regional web hosting providers never replied to any of the real abuse notifications we sent" | — | |
| | {[durumeric2014_heartbleed]} IMC 2014 | ~150,000 hosts | WHOIS abuse contact for the host IP | "a nearly 50% increase in patching by notified hosts"; 20.6% of the notified arm had begun patching within 24 h | 10.8% (not-yet-notified arm) | | | {[durumeric2014_heartbleed]} IMC 2014 | "150,000 hosts", staged in two groups | WHOIS abuse contact for the host IP | "a nearly 50% increase in patching by notified hosts"; 20.6% patching within 24 h and **39.5% after eight days**, Fisher's exact p ≈ 0 | 10.8% / **26.8%** (Group B, not yet notified) | |
| | {[stock2016_hey]} USENIX Sec 2016 | 44,790 vulnerable sites, five equal arms | RFC 2142 aliases / purchased domain WHOIS / provider abuse contacts / CERTs | of 35,832 reports sent, **2,064 (5.8%) actually received**; 74.5% still exploitable after a month; but ~40% fix //if the report is read// | 23.3% (WordPress), 2.2% (client-side XSS) | | | {[stock2016_hey]} USENIX Sec 2016 | 44,790 vulnerable sites, five equal arms | RFC 2142 aliases / purchased domain WHOIS / provider abuse contacts / CERTs | of 35,832 reports sent, **2,064 (5.8%) actually received**; 74.5% still exploitable after a month; but ~40% fix //if the report is read// | 23.3% (WordPress), 2.2% (client-side XSS) | |
| | {[li2016_youve]} USENIX Sec 2016 | 2,563 ICS + 3,536 IPv6 + 5,960 amplifier contacts | WHOIS abuse contact, national CERTs, US-CERT | "at most 18% of the population remediating" under the best regimen; direct verbose 9.8% vs national CERT 3.1% vs US-CERT 1.4% (IPv6, two days) | US-CERT arm statistically indistinguishable from control | | | {[li2016_youve]} USENIX Sec 2016 | 2,563 ICS + 3,536 IPv6 + 5,960 amplifier contacts | WHOIS abuse contact, national CERTs, US-CERT | "at most 18% of the population remediating" under the best regimen; direct verbose 9.8% vs national CERT 3.1% vs US-CERT 1.4% (IPv6, two days) | US-CERT arm statistically indistinguishable from control | |
| | {[li2016_remedying]} TheWebConf 2016 | 760,935 hijacking incidents | browser interstitial, search warning, Search Console message, WHOIS admin email | 59.5% of incidents resolved over 11 months; **82.4% / 76.8% where a Search Console alert reached a pre-registered webmaster** vs 54.6% / 43.4% otherwise | — (channel comparison) | | | {[li2016_remedying]} TheWebConf 2016 | 760,935 hijacking incidents | browser interstitial, search warning, Search Console message, WHOIS admin email | 59.5% of incidents resolved over 11 months; **82.4% / 76.8% where a Search Console alert reached a pre-registered webmaster** vs 54.6% / 43.4% otherwise | — (channel comparison) | |
| | {[stock2018_didnt]} NDSS 2018 | >24,000 domains, seven arms of ~4,000 | six email variants (plain, HTML, tracking, mailbot, S/MIME, friendly tone) | notified 24% (Git) and 17% (WordPress) fixed; **74.4% / 33.3% among those who opened the report** | 13% (Git), 14% (WordPress) | | | {[stock2018_didnt]} NDSS 2018 | >24,000 domains, seven arms of ~4,000 | six email variants (plain, HTML, tracking, mailbot, S/MIME, friendly tone) | notified 24% (Git) and 17% (WordPress) fixed; **74.4% / 33.3% among those who //viewed// the report** | 13% (Git), 14% (WordPress) | |
| | {[cetin2019_cleaning]} NDSS 2019 | ISP customers with Mirai infections | ISP walled garden vs email vs nothing | walled garden "remediates 92% of the infections within 14 days" | **74% natural remediation in the control** — the cautionary figure on this whole page | | | {[cetin2019_cleaning]} NDSS 2019 | ISP customers with Mirai infections | ISP walled garden (quarantine + notification) vs email-only vs nothing | walled garden "remediates 92% of the infections within 14 days"; **"Email-only notifications have no observable impact compared to a control group"** | **"The control group achieved the lowest cleanup rate (74%)"** with no notification at all — the cautionary figure on this whole page | |
| | {[maass2021_effective]} USENIX Sec 2021 | 4,594 German site owners, 18 arms + control | postal letter and email, addresses read by hand from each ''Impressum'' | **"56.6 % of all notified operators remediating within two months"**; 76.3% for a legal-research-group letter citing fines, 33.9% for a computer-science email citing privacy | **9.2%** | | | {[maass2021_effective]} USENIX Sec 2021 | 4,594 German site owners, 18 arms + control | postal letter and email, addresses read by hand from each ''Impressum'' | **"56.6 % of all notified operators remediating within two months"**; 76.3% for a legal-research-group letter citing fines, 33.9% for a computer-science email citing privacy | **9.2%** | |
| | {[nguyen2021_sharefirst]} USENIX Sec 2021 | 11,914 app developers | contact address from the Play Store listing | **448 responses (3.8%)** | — | | | {[nguyen2021_sharefirst]} USENIX Sec 2021 | 11,914 app developers | contact address from the Play Store listing | **448 responses (3.8%)** | — | |
| | {[squarcina2021_subdomain]} USENIX Sec 2021 | sites with dangling subdomain records | direct contact vs the authors' national CERT | direct 31% / 22% fixed vs national CERT 10% / 14% | — | | | {[squarcina2021_subdomain]} USENIX Sec 2021 | sites with dangling subdomain records | direct contact vs the authors' national CERT | direct 31% / 22% fixed vs national CERT 10% / 14% | — | |
| | {[lone2022_sav]} IEEE S&P 2022 | 2,320 network operators, RCT | 8 treatments: direct email (PeeringDB→WHOIS→abuse), national CERT/NIC.br, NOG mailing lists, × nudges | **"none of the notification treatments significantly improved SAV deployment compared to the control group"** | remediation observed in control too | | | {[lone2022_sav]} IEEE S&P 2022 | 2,320 network operators, RCT | 8 treatments: direct email (PeeringDB→WHOIS→abuse), national CERT/NIC.br, NOG mailing lists, × nudges | **"none of the notification treatments significantly improved SAV deployment compared to the control group"** | remediation observed in control too | |
| | {[sasaki2022_ics]} IEEE S&P 2022 | operators of exposed ICS devices | direct contact with the person in charge | 93 (58%) responded — "higher than most previous notification experiments" | — | | | {[sasaki2022_ics]} IEEE S&P 2022 | 160 operators of 317 exposed ICS devices | **manual telephone calls**, email only for scheduling — "the first study to directly contact the organization operating the device" | reached the person in charge for 212 devices; **"50% of the persons in charge … stated that they mitigated or will mitigate"**, and follow-up scans confirmed devices "reduced by 58% when we were able to contact the persons in charge" | 13% decrease among un-notified devices, χ² p<0.0001 | |
| | {[bennett2022_spfail]} IMC 2022 | 6,488 mail-server notifications | ''postmaster@'' per the SMTP spec | 31.6% undelivered; of 4,434 delivered, 512 (12%) opened, 177 (4%) patched, **9 (<1%) between private and public disclosure** | — | | | {[bennett2022_spfail]} IMC 2022 | 6,488 mail-server notifications | ''postmaster@'' per the SMTP spec | 31.6% undelivered; of 4,434 delivered, 512 (12%) opened, 177 (4%) patched, **9 (<1%) between private and public disclosure** | — | |
| | {[utz2023_comparing]} PoPETs 2023 | 159,035 domains, 4 privacy issues + 1 security issue | parsed privacy-policy/contact addresses vs RFC 2142 aliases | **87.8% vs 33.8% delivery**; remediation effects significant but ≤1.6 pp over control (Fisher's exact, Holm–Bonferroni) | yes, per issue | | | {[utz2023_comparing]} PoPETs 2023 | 159,035 domains, 4 privacy issues + 1 security issue | parsed privacy-policy/contact addresses vs RFC 2142 aliases | **87.8% vs 33.8% delivery**; remediation effects significant but small — the paper puts them at "0–1" to "1–2 percentage points" over control (Fisher's exact, Holm–Bonferroni), with a few later-date and generic-alias cells around 3 pp | yes, per issue | |
| | {[qin2024_rov]} NDSS 2024 | 1,012 non-deploying ASes | curated operator email rather than raw WHOIS | 824 notifications delivered, **4.07% bounce rate** — against the >50% it cites from prior work | — | | | {[qin2024_rov]} NDSS 2024 | 1,012 non-deploying ASes randomised into 6 treatment arms + control; **859 emails sent** | operator email from PeeringDB, falling back to WHOIS; nudge variants (baseline, social norms, authority, reminder, elicitation) and native language | 824 of 859 delivered, **4.07% bounce rate** against the ">50%" it cites from prior work — and **"none of the notification treatments has a significant effect"** (survival analysis; relative risk 0.46–1.35, every confidence interval spanning 1) | 11 of 138 remediated | |
| | {[elyadmani2025_keys]} IEEE S&P 2025 | 160 organisations leaking cloud-bucket secrets | leaked-file contents, OSINT, disclosure programmes, privacy-policy addresses — routed through a CSIRT partner | **95/160 (59.4%) acted**; only 20 organisations replied at all | — | | | {[elyadmani2025_keys]} IEEE S&P 2025 | 160 organisations leaking cloud-bucket secrets | leaked-file contents, OSINT, disclosure programmes, privacy-policy addresses — routed through a CSIRT partner | **95/160 (59.4%) acted**; only 20 organisations replied at all | — | |
| | {[munteanu2025_catch22]} USENIX Sec 2025 | operators of compromised hosts | Shadowserver Foundation and CERT-BUND ran the campaign | the current worked example of delegating delivery | — | | | {[munteanu2025_catch22]} USENIX Sec 2025 | operators of compromised hosts | Shadowserver Foundation and CERT-BUND ran the campaign | the current worked example of delegating delivery | — | |
| ==== The loss is in delivery, not in willingness ==== | ==== The loss is in delivery, not in willingness ==== |
| |
| The delivery funnel is the story, and it is why headline remediation rates look so bad. {[stock2016_hey]}: **5.8% of reports received**, but ~40% remediation among those actually read. {[bennett2022_spfail]}: 31.6% undelivered → 12% of the delivered opened → 4% of the openers patched. {[stock2018_didnt]}: 74.4% fixed among Git operators who opened the report, against a 13% control. Operators who read your report largely act on it. Most never read it. | The delivery funnel is the story, and it is why headline remediation rates look so bad. {[stock2016_hey]}: **5.8% of reports received**, but ~40% remediation among those actually read. {[bennett2022_spfail]}: 31.6% undelivered → 12% of the delivered opened → 4% of the openers patched. {[stock2018_didnt]}: 74.4% fixed among Git operators who //viewed// the report, against a 13% control. Operators who read your report largely act on it. Most never read it. |
| |
| So the highest-leverage thing you can do is not writing a better email. It is: | So the highest-leverage thing you can do is not writing a better email. It is: |
| ^ Factor ^ Finding ^ Source ^ | ^ Factor ^ Finding ^ Source ^ |
| | Sender identity | Legal research group beats computer-science group: 59.7% vs 54% remediation (p<0.05) | {[maass2021_effective]} | | | Sender identity | Legal research group beats computer-science group: 59.7% vs 54% remediation (p<0.05) | {[maass2021_effective]} | |
| | Framing | Legal-compliance-plus-fine beats plain GDPR beats privacy-harm; survival 50.1% / 56.6% / 69.6%, all differences significant | {[maass2021_effective]} | | | Framing | Legal-compliance-plus-fine beats plain GDPR beats privacy-harm; survival 50.1% / 56.6% / 69.6%, all differences significant. //Survival// throughout this table is the share **still non-compliant**, so lower is better | {[maass2021_effective]} | |
| | Framing | **Contradicted.** A fine warning had **no** significant effect once the notification arrived; only //receiving// it mattered | {[utz2023_comparing]} | | | Framing | **Contradicted.** A fine warning had **no** significant effect once the notification arrived; only //receiving// it mattered | {[utz2023_comparing]} | |
| | Medium | A **postal letter** beats email: survival 55.6% vs 66.3% (p<0.0001), "increasing the remediation rate by between 3.9 and 17.9 percentage points (mean: 11.1)" — for "around 5000 € on domestic postage" across 2,660 letters | {[maass2021_effective]} | | | Medium | A **postal letter** beats email: survival 55.6% vs 66.3% (p<0.0001), "increasing the remediation rate by between 3.9 and 17.9 percentage points (mean: 11.1)" — for "around 5000 € on domestic postage" across 2,660 letters | {[maass2021_effective]} | |
| * **IMC**: "Any submission that discovers security vulnerabilities must discuss their approach to responsible disclosure in their paper's Ethics section."(([[https://conferences.sigcomm.org/imc/2026/submission-instructions/|conferences.sigcomm.org/imc/2026]].)) | * **IMC**: "Any submission that discovers security vulnerabilities must discuss their approach to responsible disclosure in their paper's Ethics section."(([[https://conferences.sigcomm.org/imc/2026/submission-instructions/|conferences.sigcomm.org/imc/2026]].)) |
| * **IEEE S&P**: "Authors are required to disclose vulnerabilities no later than the rebuttal deadline. If this is not possible, the authors should notify the PC chairs by email as soon as possible."(([[https://sp2026.ieee-security.org/cfpapers.html|sp2026.ieee-security.org]].)) | * **IEEE S&P**: "Authors are required to disclose vulnerabilities no later than the rebuttal deadline. If this is not possible, the authors should notify the PC chairs by email as soon as possible."(([[https://sp2026.ieee-security.org/cfpapers.html|sp2026.ieee-security.org]].)) |
| * **USENIX Security**: submissions that fail to disclose before submission, without a convincing ethical argument for delay, may be rejected; a separate Ethical Considerations appendix is required in the final paper.((USENIX Security 2026 call for papers. ''usenix.org'' rejects automated fetches, so this was read via a search-engine cache rather than the page itself — check the live call before relying on the exact wording.)) | * **USENIX Security**: submissions that fail to disclose before submission, without a convincing ethical argument for delay, may be rejected; a separate Ethical Considerations appendix is required in the final paper.((Read from [[https://www.usenix.org/conference/usenixsecurity26/call-for-papers|usenix.org/conference/usenixsecurity26/call-for-papers]] on 2026-08-13, verbatim: "Submissions that fail to disclose prior to submission and that do not present convincing ethical arguments for delaying disclosure may be rejected" and "All papers MUST have a discussion of research ethics. This MUST be in a separate appendix called 'Ethical Considerations'". WebFetch gets 403 from usenix.org; curl with a browser User-Agent does not.)) |
| * **ACM CCS**: papers involving "real-world vulnerability analysis" must carry a dedicated Ethical Considerations section, with responsible disclosure named as an example of harm minimisation.(([[https://www.sigsac.org/ccs/CCS2026/call-for/call-for-papers.html|sigsac.org/ccs/CCS2026]].)) | * **ACM CCS**: papers involving "real-world vulnerability analysis" must carry a dedicated Ethical Considerations section, with responsible disclosure named as an example of harm minimisation.(([[https://www.sigsac.org/ccs/CCS2026/call-for/call-for-papers.html|sigsac.org/ccs/CCS2026]].)) |
| * **PoPETs** states Menlo-Report principles but no vulnerability-notification clause specifically. | * **PoPETs** states Menlo-Report principles and explicitly names "system vulnerabilities (e.g. cryptographic weaknesses, software exploits, and privacy attacks)" as an area needing ethical consideration, but has **no disclosure or notification requirement** and no timeline.((Read from [[https://petsymposium.org/authors-2026.php|petsymposium.org/authors-2026.php]] on 2026-08-13. An Ethical Principles subsection is "encouraged" and "may be required if deemed necessary during the review process".)) |
| |
| The practical consequence: **the notification has to happen before you submit, so it belongs in your timeline from the start.** A campaign that needs a month of monitoring plus a reminder schedule cannot be bolted on in the week before a deadline, and "we will disclose after acceptance" is now grounds for rejection at two of these venues. | The practical consequence: **the notification has to happen before you submit, so it belongs in your timeline from the start.** A campaign that needs a month of monitoring plus a reminder schedule cannot be bolted on in the week before a deadline, and "we will disclose after acceptance" is now grounds for rejection at two of these venues. |
| |
| * **Seven venues only.** EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS are absent, and CHI/SOUPS are where a good deal of the operator-facing usable-security work appears. NDSS 2016 and NDSS 2018 full text was not retrieved at all, which is why {[stock2018_didnt]} — the canonical follow-up study — was read from the publisher's PDF rather than from the corpus. | * **Seven venues only.** EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS are absent, and CHI/SOUPS are where a good deal of the operator-facing usable-security work appears. NDSS 2016 and NDSS 2018 full text was not retrieved at all, which is why {[stock2018_didnt]} — the canonical follow-up study — was read from the publisher's PDF rather than from the corpus. |
| * ''ethics.notifiedAffectedParties'' **agreed with an independent re-extraction on 67% of papers**, so treat the percentages as accurate to a few points, not to the decimal. ''ethics.disclosureDetail'' is free text capped at 20 words and is reported below only as folded families with the residue printed. | * ''ethics.notifiedAffectedParties'' **agreed with an independent re-extraction on 67% of papers**, so treat every percentage above as accurate to a few points, not to the decimal.((That 67% was measured on 100 papers of the previous, 4,322-paper extraction run and has not been re-measured on the current corpus. Treat it as the right order of magnitude.)) ''ethics.disclosureDetail'' is free text capped at 20 words, so it is reported only as folded families and only as a ranking; the fold rules and the unmapped residue are on the provenance page. |
| * **The 22 campaigns are hand-classified, not a census.** A regex over 5,869 full texts produced 179 candidates; 22 were classified as campaigns, 16 rejected with a stated reason, and **144 were never read**. The table above is a curated reading list, and the unreviewed residue is published in full. | * **The 22 campaigns are hand-classified, not a census.** A regex over 5,869 full-text files (ten more than the 5,859 extraction records — a few papers have text but no record) produced 179 candidates; 22 were classified as campaigns, 13 rejected with a stated reason, and **144 were never read**. Three further rejections came from a wider first-pass scan and fall outside the committed regex, which is why the reasoned rejections total 16. The table above is a curated reading list, and the unreviewed residue is published in full. |
| |
| The complete query log, the report script with its unedited output, the folding rules with their unmapped residue, the spot-checked quotes, and every external source that was rejected are on **[[provenance:practices:notifying_websites]]**. Corpus-wide caveats are on [[Literature:Corpus]]. | The complete query log, the report script with its unedited output, the folding rules with their unmapped residue, the spot-checked quotes, and every external source that was rejected are on **[[provenance:practices:notifying_websites]]**. Corpus-wide caveats are on [[Literature:Corpus]]. |