| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| design:website_classification [2026/08/12 10:05] – Review pass (Fable): fix stale figures that sat OUTSIDE the corpus section and so were missed by the windowed staleness guard. Lead paragraph 72% -> 75% for bespoke unsized taxonomies. Authored by Claude. karel.kubicek.claude | design:website_classification [2026/08/27 14:39] (current) – CrUX licence: Tranco homepage CC BY-SA 4.0, methodology silent, Google CC BY 4.0. Authored by Claude karel.kubicek.claude |
|---|
| === VirusTotal === | === VirusTotal === |
| |
| * **API**: ''GET /api/v3/domains/{domain}'' returns a ''categories'' attribute, //"a dictionary mapping that relates categorisation services with the category it assigns the domain to"//. Documented at [[https://docs.virustotal.com/reference/domains-object|docs.virustotal.com]]. | * **API**: ''GET /api/v3/domains/{domain}'' returns a ''categories'' attribute, //"a dictionary mapping that relates categorisation services with the category it assigns the domain to"//. Documented at [[https://docs.virustotal.com/reference/domains-object|docs.virustotal.com]]. **That is a topic label, not a maliciousness verdict.** The oracle use of the same API — engine disagreement, thresholds, URL versus file lookup, snapshot dating — is [[security:virustotal]]. |
| * **Which vendors**: Dr.Web, Forcepoint ThreatSeeker, BitDefender, Sophos, Trend Micro, Websense and (legacy) Alexa, among others — so several of the vendors above reach you through VirusTotal without you querying them. | * **Which vendors**: Dr.Web, Forcepoint ThreatSeeker, BitDefender, Sophos, Trend Micro, Websense and (legacy) Alexa, among others — so several of the vendors above reach you through VirusTotal without you querying them. |
| * **Rate limit**: the free tier is **500 requests per day and 4 requests per minute**.((https://docs.virustotal.com/docs/public-vs-premium-api — fetched 2026-08-07; that page names the "Community" tier but does not mention the October 2025 retiering, so treat the tier names as unconfirmed and the numbers as current.)) The same page adds that the Public API //"must not be used in commercial products or services"//. At 500/day, a 100k-domain study takes 200 days. Vallina et al. had an academic key at 20k/day; we could find no self-serve academic application form in 2026, so budget for a direct conversation with VirusTotal rather than assuming access. | * **Rate limit**: the free tier is **500 requests per day and 4 requests per minute**.((https://docs.virustotal.com/docs/public-vs-premium-api — fetched 2026-08-07; that page names the "Community" tier but does not mention the October 2025 retiering, so treat the tier names as unconfirmed and the numbers as current.)) The same page adds that the Public API //"must not be used in commercial products or services"//. At 500/day, a 100k-domain study takes 200 days. Vallina et al. had an academic key at 20k/day. As of 2026-08-27 there is still no standalone quota-application form; the public contact form has a subject "I have an academic research request" ([[https://www.virustotal.com/gui/contact-us/legal|virustotal.com/gui/contact-us/legal]]). Budget for that conversation rather than assuming 20k/day. Maliciousness-oracle use of the same API is [[security:virustotal]]. |
| * **Advantages**: aggregates many providers in a single call; widely used and easy to cite. | * **Advantages**: aggregates many providers in a single call; widely used and easy to cite. |
| * **Disadvantages**: **the integration is lossy**. In the 2020 audit, several services returned labels when queried directly but not through VirusTotal, and Trend Micro's VirusTotal labels tracked its 2011 taxonomy rather than its 2019 one. You do not control which product version you are reading, and "we used VirusTotal categories" does not identify the underlying source. | * **Disadvantages**: **the integration is lossy**. In the 2020 audit, several services returned labels when queried directly but not through VirusTotal, and Trend Micro's VirusTotal labels tracked its 2011 taxonomy rather than its 2019 one. You do not control which product version you are reading, and "we used VirusTotal categories" does not identify the underlying source. |
| === Alexa === | === Alexa === |
| |
| Amazon retired **Alexa.com on 1 May 2022**, per its own end-of-service notice ("we will be retiring Alexa.com on May 1, 2022").((Captured on the alexa.com login page, https://web.archive.org/web/20220315000000/https://www.alexa.com/ — retrieved 2026-08-07. [[Design:Website Selection]] currently gives 1 August 2023; we could not find a primary source for that date, and ''alexa.com'' now redirects to the unrelated Amazon Alexa voice assistant.)) Its ranking service and its category service died together. | Amazon retired **Alexa.com on 1 May 2022**, per its own end-of-service notice ("we will be retiring Alexa.com on May 1, 2022").((Captured on the alexa.com login page, https://web.archive.org/web/20220315000000/https://www.alexa.com/ — retrieved 2026-08-07. An earlier revision of [[Design:Website Selection]] gave 1 August 2023; that date is when Tranco dropped Alexa from the default list, not when Amazon switched the service off. The selection page now carries both dates. ''alexa.com'' redirects to the unrelated Amazon Alexa voice assistant.)) Its ranking service and its category service died together. |
| |
| * **Why it still matters**: it appears in 12 of the 330 corpus papers that categorise websites, under eight different spellings, and papers published as late as 2024 still use it because their data collection predates the shutdown. If you are reading such a paper, the labels are from a dead service with a documented 0.53% coverage rate. | * **Why it still matters**: it appears in 12 of the 330 corpus papers that categorise websites, under eight different spellings, and papers published as late as 2024 still use it because their data collection predates the shutdown. If you are reading such a paper, the labels are from a dead service with a documented 0.53% coverage rate. |
| |
| * **The TLD.** Precise where it exists and absent where it matters — a ccTLD is strong evidence, but generic TLDs carry no country at all, and that is most of the head of any list. | * **The TLD.** Precise where it exists and absent where it matters — a ccTLD is strong evidence, but generic TLDs carry no country at all, and that is most of the head of any list. |
| * **CrUX country lists, via Tranco.** Tranco's list-generation API takes ''filterCRUX'', ''filterCRUXType'' (''global'' / ''country'' / ''region'' / ''subregion''), ''filterCRUXValue'' (e.g. a list of country codes) and ''filterCRUXMonth'', so you can generate a country-restricted list reproducibly, with a permalink, from a free account.((https://tranco-list.eu/api_documentation — fetched 2026-08-07. Basic Auth with your email as username and API token as password; the ''/configure'' page returns 401 without a login. CrUX itself is CC BY-SA 4.0.)) This is the cleanest free source and it is under-used. **But it ranks by page loads from Chrome users in a country, which is popularity, not audience** — and the overlap is severe: the union of five country top-10k lists (China, Germany, Italy, Korea, Turkey) is only 18,718 domains rather than 50,000, with 4,017 domains common to all five, and a preliminary labelling built from these lists is incompatible with the site's own TLD in at least 25% of cases for every label {[bozzolan2026_llmweb]}. | * **CrUX country lists, via Tranco.** Tranco's list-generation API takes ''filterCRUX'', ''filterCRUXType'' (''global'' / ''country'' / ''region'' / ''subregion''), ''filterCRUXValue'' (e.g. a list of country codes) and ''filterCRUXMonth'', so you can generate a country-restricted list reproducibly, with a permalink, from a free account.((https://tranco-list.eu/api_documentation — fetched 2026-08-07. Basic Auth with your email as username and API token as password; the ''/configure'' page returns 401 without a login. CrUX datasets are CC BY 4.0 per Google's methodology, fetched 2026-08-27; Tranco's homepage still labels CrUX CC BY-SA 4.0, while the methodology page states no CrUX licence.)) This is the cleanest free source and it is under-used. **But it ranks by page loads from Chrome users in a country, which is popularity, not audience** — and the overlap is severe: the union of five country top-10k lists (China, Germany, Italy, Korea, Turkey) is only 18,718 domains rather than 50,000, with 4,017 domains common to all five, and a preliminary labelling built from these lists is incompatible with the site's own TLD in at least 25% of cases for every label {[bozzolan2026_llmweb]}. |
| * **Site language.** A good proxy, and cheap, but it splits badly on English, Spanish, Arabic and Portuguese, which is a large share of the web. | * **Site language.** A good proxy, and cheap, but it splits badly on English, Spanish, Arabic and Portuguese, which is a large share of the web. |
| * **Host IP or CDN location.** Measures where bytes are served from. Behind a CDN — most of the head of any list — it tells you about the CDN. See [[Design:IP Classification]] and [[Design:Crawling Location]]. | * **Host IP or CDN location.** Measures where bytes are served from. Behind a CDN — most of the head of any list — it tells you about the CDN. See [[Design:IP Classification]] and [[Design:Crawling Location]]. |
| - Based on LinkedIn profiles that are self-reported — prone to adversarial data. | - Based on LinkedIn profiles that are self-reported — prone to adversarial data. |
| - Only a subset of PeopleDataLabs' full dataset. The "22M of 70M rows" framing is longstanding on this page; the 22M is confirmed, the 70M total could not be re-confirmed in 2026 and PDL's marketing page now cites 23.8M+ without saying which corpus that is. | - Only a subset of PeopleDataLabs' full dataset. The "22M of 70M rows" framing is longstanding on this page; the 22M is confirmed, the 70M total could not be re-confirmed in 2026 and PDL's marketing page now cites 23.8M+ without saying which corpus that is. |
| <wrap todo>TODO: cite ''Machine Learning Compliance Analysis for Email Regulation'' when it is public.</wrap> | <WRAP todo>TODO: cite ''Machine Learning Compliance Analysis for Email Regulation'' when it is public.</WRAP> |
| |
| ==== Crunchbase ==== | ==== Crunchbase ==== |
| - Focuses mostly on variables useful for investments and market competitiveness. | - Focuses mostly on variables useful for investments and market competitiveness. |
| - Academic access is no longer publicly documented; budget for a sales conversation. | - Academic access is no longer publicly documented; budget for a sales conversation. |
| <wrap todo>TODO: cite ''Machine Learning Compliance Analysis for Email Regulation'' when it is public.</wrap> | <WRAP todo>TODO: cite ''Machine Learning Compliance Analysis for Email Regulation'' when it is public.</WRAP> |
| |
| ==== Orbis ==== | ==== Orbis ==== |
| * **Enum fields versus free text.** Method and validation are enums, stable enough to publish as rough shares (''classification.method'' agrees 58% run-to-run, so read those as a ranking). Service names and taxonomies are free text and are reported as rankings and folded families only. | * **Enum fields versus free text.** Method and validation are enums, stable enough to publish as rough shares (''classification.method'' agrees 58% run-to-run, so read those as a ranking). Service names and taxonomies are free text and are reported as rankings and folded families only. |
| * **Quotes were checked.** Every figure above traces to tuples carrying a verbatim evidence quote; a sample of these was re-located in the source PDFs. Of six quotes checked by hand, two initially "failed" a literal grep and turned out to be intact but split across a two-column break — normalise whitespace before concluding that a quote is not in the paper. The five new LLM website-category tuples were re-checked individually on 2026-08-12; see [[provenance:design:website_classification]]. | * **Quotes were checked.** Every figure above traces to tuples carrying a verbatim evidence quote; a sample of these was re-located in the source PDFs. Of six quotes checked by hand, two initially "failed" a literal grep and turned out to be intact but split across a two-column break — normalise whitespace before concluding that a quote is not in the paper. The five new LLM website-category tuples were re-checked individually on 2026-08-12; see [[provenance:design:website_classification]]. |
| * **Coverage.** EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS are absent entirely. The 2025 and 2026 venue-years are incomplete by construction rather than by relevance, so any row that reaches them is a floor. Notably, **{[vallina2020_misshapes]} itself is in the venue index but has no extracted full text** — the reference work for this page is not in the population the page measures. | * **Coverage.** Seven venues only, with 2025 and 2026 incomplete by construction rather than by relevance, so any row that reaches them is a floor; the scope and the selection funnel are on [[literature:corpus]]. Notably, **{[vallina2020_misshapes]} itself is in the venue index but has no extracted full text** — the reference work for this page is not in the population the page measures. |
| | * **Every query behind this section, the report script and its unedited output** are on [[provenance:design:website_classification]]; corpus-level caveats are on [[literature:corpus]]. |
| |
| ===== Open Questions ===== | ===== Open Questions ===== |