User Tools

Site Tools


design:ip_classification

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
design:ip_classification [2026/08/14 03:20] – programming:traffic_files now exists — drop the 'not yet written' marker. Authored by Claude. karel.kubicek.claudedesign:ip_classification [2026/08/21 14:50] (current) – Boxes: <wrap> renders a span, use uppercase <WRAP>; one box per list, not per bullet karel.kubicek.claude
Line 72: Line 72:
   * **[[https://asdb.stanford.edu/|ASdb]]** {[ziv2021_asdb]} classifies ASNs into 17 industry categories and 95 sub-categories (hosting, ISP, education, government…), reporting 96% coverage of ASes at 93% accuracy on the top level and 75% on sub-categories. Snapshots run through March 2026. Used, for example, by {[darwich2023_replication]} to describe what kind of networks their geolocation targets sat in.   * **[[https://asdb.stanford.edu/|ASdb]]** {[ziv2021_asdb]} classifies ASNs into 17 industry categories and 95 sub-categories (hosting, ISP, education, government…), reporting 96% coverage of ASes at 93% accuracy on the top level and 75% on sub-categories. Snapshots run through March 2026. Used, for example, by {[darwich2023_replication]} to describe what kind of networks their geolocation targets sat in.
  
-<wrap todo>**CAIDA's separate //AS Classification// dataset is discontinued** — CAIDA's own page states it is no longer supported and download access has been removed. Papers from 2015–2022 cite it routinely; ASdb is the live replacement. Check before you cite a dataset you found in a related-work section.</wrap>+<WRAP todo>**CAIDA's separate //AS Classification// dataset is discontinued** — CAIDA's own page states it is no longer supported and download access has been removed. Papers from 2015–2022 cite it routinely; ASdb is the live replacement. Check before you cite a dataset you found in a related-work section.</WRAP>
  
 The current state of the art on the sibling problem is {[selmo2025_borges]} (IMC 2025), which few-shot prompts an LLM over PeeringDB free-text fields plus website and domain evidence, and reports a 7% improvement in sibling-ASN identification over AS2Org-style methods. That is, as of 2026, **the one place in IP classification where an LLM method has cleared peer review** — see [[#Open Questions]]. The current state of the art on the sibling problem is {[selmo2025_borges]} (IMC 2025), which few-shot prompts an LLM over PeeringDB free-text fields plus website and domain evidence, and reports a 7% improvement in sibling-ASN identification over AS2Org-style methods. That is, as of 2026, **the one place in IP classification where an LLM method has cleared peer review** — see [[#Open Questions]].
Line 233: Line 233:
 **IPv6 is different, and in the direction people do not expect.** The intuition is that IPv6 clients rotate temporary addresses ([[https://www.rfc-editor.org/rfc/rfc8981.html|RFC 8981]]) and are therefore harder to track. That is true of the interface identifier and false of the assignment: {[padmanabhan2020_dynamips]} found that **IPv6 assignments last //longer// than IPv4 ones, often remaining stable for months** — which makes long-term tracking of an IPv6 subscriber //easier//, not harder.((Paraphrased rather than quoted: we verified the paper, venue and topic against Crossref and the authors' own listing, but could not reach a copy to quote the sentence verbatim. Check it before citing the wording.)) Aggregate on the /64 or the delegated prefix, not the full address, and do not assume dual-stack clients are equally identifiable on both stacks. Where hardware still derives the interface identifier from a MAC address, the address is //more// identifying than IPv4 ever was: {[rye2023_ipvseeyou]} extracted EUI-64-derived MACs from over 12M routers in 146 countries and geolocated them by correlating with wardriving data, reporting a **median error of 39 metres**. For the same devices, MaxMind's locations sat a median **26 km** from those wardriving positions — three orders of magnitude apart, from the same input address. **IPv6 is different, and in the direction people do not expect.** The intuition is that IPv6 clients rotate temporary addresses ([[https://www.rfc-editor.org/rfc/rfc8981.html|RFC 8981]]) and are therefore harder to track. That is true of the interface identifier and false of the assignment: {[padmanabhan2020_dynamips]} found that **IPv6 assignments last //longer// than IPv4 ones, often remaining stable for months** — which makes long-term tracking of an IPv6 subscriber //easier//, not harder.((Paraphrased rather than quoted: we verified the paper, venue and topic against Crossref and the authors' own listing, but could not reach a copy to quote the sentence verbatim. Check it before citing the wording.)) Aggregate on the /64 or the delegated prefix, not the full address, and do not assume dual-stack clients are equally identifiable on both stacks. Where hardware still derives the interface identifier from a MAC address, the address is //more// identifying than IPv4 ever was: {[rye2023_ipvseeyou]} extracted EUI-64-derived MACs from over 12M routers in 146 countries and geolocated them by correlating with wardriving data, reporting a **median error of 39 metres**. For the same devices, MaxMind's locations sat a median **26 km** from those wardriving positions — three orders of magnitude apart, from the same input address.
  
-<wrap todo>Before you use an IP as a user identifier, write down which of these four you have ruled out and how. If the answer is "none", use it as a network identifier (prefix or ASN) instead, where all four are far weaker.</wrap>+<WRAP todo>Before you use an IP as a user identifier, write down which of these four you have ruled out and how. If the answer is "none", use it as a network identifier (prefix or ASN) instead, where all four are far weaker.</WRAP>
  
 ===== A Script ===== ===== A Script =====
Line 707: Line 707:
 ===== Open Questions ===== ===== Open Questions =====
  
-  * <wrap todo>**No head-to-head of the geolocation databases on a web-measurement population.** {[gharaibeh2017_look]} used router interfaces (2017), {[darwich2023_replication]} RIPE Atlas anchors (2023), {[nabi2026_lostprefix]} Atlas and Giga (2026). None of them is "the servers a Tranco top-10k crawl connects to" — a population dominated by CDN and cloud edges, which is exactly the population the anycast and datacenter caveats bite hardest on.</wrap> +<WRAP todo> 
-  * <wrap todo>**How much does the choice of database move a published cross-border-transfer figure?** Re-running one compliance paper's analysis with four databases would be a small, cheap, useful replication, and the answer is not obviously small: the run above disagreed on country for the AWS edge.</wrap> +  * **No head-to-head of the geolocation databases on a web-measurement population.** {[gharaibeh2017_look]} used router interfaces (2017), {[darwich2023_replication]} RIPE Atlas anchors (2023), {[nabi2026_lostprefix]} Atlas and Giga (2026). None of them is "the servers a Tranco top-10k crawl connects to" — a population dominated by CDN and cloud edges, which is exactly the population the anycast and datacenter caveats bite hardest on. 
-  * <wrap todo>**Geofeed adoption is not measured for the web.** {[livadariu2024_geofeeds]} gives 1.50% of allocated IPv4 prefixes overall; nobody has asked what share of the //address space a crawl actually touches// has a geofeed, which — given how concentrated that space is on a few large operators — could be much higher or much lower.</wrap> +  * **How much does the choice of database move a published cross-border-transfer figure?** Re-running one compliance paper's analysis with four databases would be a small, cheap, useful replication, and the answer is not obviously small: the run above disagreed on country for the AWS edge. 
-  * <wrap todo>**No systematic evaluation of the commercial datacenter/VPN/proxy flags against ground truth.** Every paper that uses one takes it on faith, and the run above shows two public DNS resolvers flagged ''is_vpn'' and ''is_abuser''. A labelled benchmark here would be immediately useful and is well within a single student's reach.</wrap> +  * **Geofeed adoption is not measured for the web.** {[livadariu2024_geofeeds]} gives 1.50% of allocated IPv4 prefixes overall; nobody has asked what share of the //address space a crawl actually touches// has a geofeed, which — given how concentrated that space is on a few large operators — could be much higher or much lower. 
-  * <wrap todo>**No peer-reviewed method for identifying hosting/datacenter address space** beyond the operators' own published lists — which means everything outside the big five clouds is guesswork.</wrap> +  * **No systematic evaluation of the commercial datacenter/VPN/proxy flags against ground truth.** Every paper that uses one takes it on faith, and the run above shows two public DNS resolvers flagged ''is_vpn'' and ''is_abuser''. A labelled benchmark here would be immediately useful and is well within a single student's reach. 
-  * <wrap todo>**LLMs have reached AS-to-organisation mapping {[selmo2025_borges]} but not IP classification.** We found nothing peer-reviewed applying an LLM to geolocation, host typing or residential/VPN/datacenter labelling as of August 2026 — unlike cookie and policy classification, where LLM methods are now routine. Whether that is because the task has no useful text to read, or because nobody has tried, is an open question.</wrap> +  * **No peer-reviewed method for identifying hosting/datacenter address space** beyond the operators' own published lists — which means everything outside the big five clouds is guesswork. 
-  * <wrap todo>**No strong successor to {[richter2016_multi]} on CGNAT prevalence.** The best numbers on how much of the client Internet sits behind shared addresses are a decade old, and IPv4 exhaustion has only got worse since.</wrap>+  * **LLMs have reached AS-to-organisation mapping {[selmo2025_borges]} but not IP classification.** We found nothing peer-reviewed applying an LLM to geolocation, host typing or residential/VPN/datacenter labelling as of August 2026 — unlike cookie and policy classification, where LLM methods are now routine. Whether that is because the task has no useful text to read, or because nobody has tried, is an open question. 
 +  * **No strong successor to {[richter2016_multi]} on CGNAT prevalence.** The best numbers on how much of the client Internet sits behind shared addresses are a decade old, and IPv4 exhaustion has only got worse since. 
 +</WRAP>
  
 ===== Related Pages ===== ===== Related Pages =====
design/ip_classification.1786677608.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki