programming:registration
Differences
This shows you the differences between two versions of the page.
| Both sides previous revisionPrevious revisionNext revision | Previous revision | ||
| programming:registration [2025/02/18 15:27] – formatting fix karelkubicek | programming:registration [2026/08/27 13:08] (current) – Apply focused+generic review: split Cookie Hunter 13.7% vs 25,242; Kubicek 25.7% of unique domains; login_diff segment match and aria-label; recaptcha deprecation; no ToS-consensus claim. Authored by Claude karel.kubicek.claude | ||
|---|---|---|---|
| Line 1: | Line 1: | ||
| ====== Automating Login and Registration to Websites ====== | ====== Automating Login and Registration to Websites ====== | ||
| - | WiP: brainstorming below. | + | Logged-out is the default crawl in this literature. Of the 857 web crawls in this corpus, **553 (64.5%)** are labelled '' |
| - | ===== Automated Login ===== | + | <WRAP important> |
| + | Do not read the 81, or the 42 + 24 + 15 split next to it, as " | ||
| + | </ | ||
| - | Shepherd by Jonker et al. (2020) | + | Getting past a login is a different job from clicking a subpage. Depth, typing-versus-'' |
| - | ===== SSO ===== | + | ===== What to read first ===== |
| - | How to use it for large scale web measurements, references | + | ^ Paper ^ Why ^ |
| + | | Drakonakis et al., CCS 2020, //Cookie Hunter// {[drakonakis2020_cookie]} | The scale-registration baseline: **13.7%** registered and logged in of 168,594 signup domains; **25,242** accounts created (not the same figure). They **refused** a human CAPTCHA farm | | ||
| + | | Rautenstrauch et al., IEEE S&P 2024, //To Auth or Not To Auth// {[rautenstrauch2024_auth]} | Why you bother: 200 sites where automated login then kept working; more than 400 accounts made by hand; the failure modes that ate the rest of the list | | ||
| + | | Kaizer et al., IMC 2016 {[kaizer2016_characterizing]} | The existence proof that logged-in and logged-out are different websites: 345 sites, 14 Alexa categories, accounts created by hand, Selenium for the login | | ||
| + | | Kubicek et al., TheWebConf 2024 {[kubicek2024_register]} | The largest registration crawl in the bibliographic index. **In the index, absent from the extraction** — cited from the author PDF, not from '' | ||
| + | | Jonker et al., MADWeb 2020 {[jonker2020_shepherd]} | How to //verify// a login. Out of this corpus (NDSS workshop). 7,113 verified logins with BugMeNot credentials | | ||
| + | | Englehardt et al., PoPETs 2018 {[englehardt2018_email]} | Newsletter signup is a lighter instrument than full registration: 3,335 form attempts, 12,618 emails from 902 senders | | ||
| - | * Dimova et al. (2023) | + | ===== Four jobs ===== |
| - | * Ghasamisharif et al. (2018) | + | |
| - | * Zhou et al. (2014) - might be too old | + | |
| - | * any other? | + | |
| - | ===== Automated Registration | + | - **Login with existing credentials.** Shepherd {[jonker2020_shepherd]} is the method paper. Calzavara et al. {[calzavara2014_cookiejar]} is the cookie-identification paper (70 sites). Kaizer et al. {[kaizer2016_characterizing]} |
| + | - **Create accounts at scale.** Cookie Hunter {[drakonakis2020_cookie]}, | ||
| + | - **Newsletter | ||
| + | - **Fill without submitting.** Senol et al. {[senol2022_leaky]} filled email on 52,055 of 99,380 loaded Tranco 100k sites and left. Chatzimpyrros et al. {[chatzimpyrros2020_register]} did the registration-form version on 200,000 sites, **outside this corpus**. Typing versus setting '' | ||
| - | * Drakonakis et al. (2020) | + | ===== What the field actually did ===== |
| - | * Kubicek et al. (2024) | + | |
| - | * Englehardt et al. (2018) and Mathur et al. (2020) follow-up | + | |
| + | Population: **857 web crawls** ('' | ||
| - | ===== Other ===== | + | ^ '' |
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | (no '' | ||
| + | | '' | ||
| + | | '' | ||
| + | The 42 + 24 + 15 = 81. Hand-read, they split like this. A paper can do more than one of these; the label is the job that decided whether it belongs on this page. | ||
| - | * Automated interaction wit hforms but without submission: Senol et al. (2022) and Chatzimpyrros et al. (2019) | + | ^ Role ^ Papers ^ Share of 81 ^ What it is ^ |
| + | | scale-register | 8 | 9.9% | automated account creation on the sites being measured | | ||
| + | | scale-login | 9 | 11.1% | automated login with credentials that already exist | | ||
| + | | **instrument | ||
| + | | manual-session | 9 | 11.1% | a human created the account or logged in; the crawler reused it | | ||
| + | | lab-scanner | 14 | 17.3% | authenticated scan of a web application they control | | ||
| + | | sso-study | 2 | 2.5% | SSO is the phenomenon, not a way into a ranked list | | ||
| + | | newsletter | 1 | 1.2% | mailing-list / marketing-email registration inside the 81 | | ||
| + | | marketplace | 6 | 7.4% | underground / forum / darkweb session | | ||
| + | | social | 10 | 12.3% | accounts on one platform | ||
| + | | other | 11 | 13.6% | real accounts, not a ranked-list login crawl | | ||
| + | | mislabel | 11 | 13.6% | did not authenticate to the sites they measured | | ||
| + | Four papers the schema labelled '' | ||
| - | /* | + | Kubicek et al. {[kubicek2024_register]} |
| - | This is a comment not visible on the page. It outlines the syntax | + | |
| - | ===== Header Level 2 ===== | + | The 17 instrument papers, in year order: SSOScan {[zhou2014_ssoscan]}; |
| - | ==== Header Level 3 ==== | + | |
| - | === Links === | + | Year shape of the 17, of web crawls in that bucket. 2025–2026 is provisional: |
| - | External links are recognized automatically: | + | ^ Bucket ^ Web crawls ^ Schema-logged ^ Instrument ^ |
| + | | 2010–2013 | 80 | 7 | 0 | | ||
| + | | 2014–2017 | 130 | 18 | 6 | | ||
| + | | 2018–2021 | 241 | 14 | 2 | | ||
| + | | 2022–2024 | 253 | 25 | 7 | | ||
| + | | 2025–2026* | 153 | 17 | 2 | | ||
| - | Internal links are created by using square brackets. You can either just give a [[pagename]] or use an additional [[pagename|link text]]. | + | ===== Automated login ===== |
| - | === Lists === | + | The method that is still the one to copy is Shepherd {[jonker2020_shepherd]}. **Out of this corpus** (MADWeb 2020, an NDSS workshop). Four steps: find a login start, submit, check the response, **verify**. The last step is the one people skip. Shepherd runs a check twice — once with the candidate session' |
| - | Lists and their levels are decided by indentation | + | On BugMeNot credentials for 49,846 sites, 23,088 were rejected as invalid, 26,758 remained, |
| - | - this is 1. item | + | Calzavara et al. {[calzavara2014_cookiejar]} |
| - | - 2. item | + | |
| - | - nested a. item | + | |
| - | * bullet-point item | + | |
| - | === Code === | + | Kaizer et al. {[kaizer2016_characterizing]} created accounts **manually** on 345 sites in 14 Alexa categories, then used Selenium to log in. Logged-in users saw more ads and more privacy-concerning behaviour. If your question is "what does the logged-in web look like", this is still the paper that posed it cleanly. |
| - | For a short inline monospace, use '' | + | Rautenstrauch et al. {[rautenstrauch2024_auth]} is the 2024 version of that question for security headers and XSS. Semi-automatic account framework; **more than 400 accounts by hand**; then **200 sites** where automated login kept succeeding. Playwright 1.33, Gmail for email verification. The rest of the list died of captcha loops / timeouts / blocks |
| - | <code python> | + | The script below is Shepherd' |
| - | string = "World" | + | |
| - | print(f' | + | |
| - | </ | + | |
| - | For large code, use '' | + | < |
| + | # | ||
| + | """ | ||
| - | <file php example.php> | + | A 200 and a page that still has a login form is not a login. Jonker, Karsch, |
| - | <?php echo "hello world!"; | + | Krumnow and Sleegers (MADWeb 2020) verify a candidate login by running the |
| + | same check twice: once with the cookies from the candidate session, once | ||
| + | without. A login is claimed only when the check succeeds with cookies and | ||
| + | fails without them. | ||
| + | |||
| + | This script is that comparison on two HTML snapshots. It does not find the | ||
| + | login form, submit credentials, | ||
| + | Shepherd steps; this is the one that is easy to get wrong by substring. | ||
| + | |||
| + | python3 login_diff.py | ||
| + | python3 login_diff.py --logged-out out.html --logged-in in.html | ||
| + | |||
| + | Exit 0 if the logged-in snapshot has at least one login marker that the | ||
| + | logged-out snapshot does not. Exit 1 otherwise, or if a self-test fails. | ||
| + | """ | ||
| + | from __future__ import annotations | ||
| + | |||
| + | import argparse | ||
| + | import re | ||
| + | import sys | ||
| + | |||
| + | # Attribute values are split into path segments. A segment matches a keyword | ||
| + | # only if it equals the keyword, or the keyword plus a button/link suffix. | ||
| + | # So " | ||
| + | # do not. Substring search of the HTML is how those three false-positived. | ||
| + | ATTR_RE = re.compile( | ||
| + | r" | ||
| + | re.IGNORECASE, | ||
| + | ) | ||
| + | SUFFIXES = ("", | ||
| + | LOGOUT_KW = (" | ||
| + | USER_KW = (" | ||
| + | LOGIN_KW = (" | ||
| + | |||
| + | |||
| + | def segments(value: | ||
| + | parts = re.split(r" | ||
| + | out = {p.strip() for p in parts if p.strip()} | ||
| + | # aria-label=" | ||
| + | # to hyphen so it matches the keyword " | ||
| + | # would turn the label into the tokens " | ||
| + | out.add(value.lower().replace(" | ||
| + | return out | ||
| + | |||
| + | |||
| + | def hits(html: str, keywords: tuple[str, ...]) -> set[str]: | ||
| + | found: set[str] = set() | ||
| + | allowed = {kw + suf for kw in keywords for suf in SUFFIXES} | ||
| + | for m in ATTR_RE.finditer(html): | ||
| + | for seg in segments(m.group(1)): | ||
| + | if seg in allowed: | ||
| + | found.add(seg) | ||
| + | return found | ||
| + | |||
| + | |||
| + | def markers(html: | ||
| + | return { | ||
| + | " | ||
| + | " | ||
| + | " | ||
| + | } | ||
| + | |||
| + | |||
| + | LOGGED_OUT_HTML = """ | ||
| + | < | ||
| + | <a href="/ | ||
| + | <form id=" | ||
| + | <input type=" | ||
| + | <button type=" | ||
| + | </ | ||
| + | < | ||
| + | </ | ||
| + | """ | ||
| + | |||
| + | LOGGED_IN_HTML = """< | ||
| + | < | ||
| + | <a href="/ | ||
| + | <a href="/ | ||
| + | <div data-testid=" | ||
| + | </ | ||
| + | """ | ||
| + | |||
| + | TRAP_HTML = """< | ||
| + | < | ||
| + | < | ||
| + | < | ||
| + | <a href="/ | ||
| + | </ | ||
| + | """ | ||
| + | |||
| + | |||
| + | def verdict(logged_out: | ||
| + | """ | ||
| + | out_m = markers(logged_out) | ||
| + | in_m = markers(logged_in) | ||
| + | reasons: list[str] = [] | ||
| + | |||
| + | logout_only_in = in_m[" | ||
| + | if logout_only_in: | ||
| + | reasons.append(" | ||
| + | |||
| + | user_only_in = in_m[" | ||
| + | if user_only_in: | ||
| + | reasons.append(" | ||
| + | |||
| + | login_only_out = out_m[" | ||
| + | if out_m[" | ||
| + | reasons.append( | ||
| + | "login area present logged-out and absent logged-in: " + ", " | ||
| + | ) | ||
| + | |||
| + | return bool(reasons), | ||
| + | |||
| + | |||
| + | def self_test() -> None: | ||
| + | failed = 0 | ||
| + | |||
| + | ok, why = verdict(LOGGED_OUT_HTML, | ||
| + | if not ok: | ||
| + | print(" | ||
| + | failed += 1 | ||
| + | else: | ||
| + | print(" | ||
| + | for line in why: | ||
| + | print(" | ||
| + | |||
| + | ok, _why = verdict(LOGGED_OUT_HTML, | ||
| + | if ok: | ||
| + | print(" | ||
| + | failed += 1 | ||
| + | else: | ||
| + | print(" | ||
| + | |||
| + | ok, _why = verdict(TRAP_HTML, | ||
| + | if ok: | ||
| + | print(" | ||
| + | failed += 1 | ||
| + | else: | ||
| + | print(" | ||
| + | |||
| + | # Substring traps the first regex version accepted. Both snapshots keep the | ||
| + | # login form; the candidate only adds the tempting attribute. | ||
| + | for name, extra in ( | ||
| + | (" | ||
| + | (" | ||
| + | (" | ||
| + | ): | ||
| + | candidate = LOGGED_OUT_HTML.replace("</ | ||
| + | ok, why = verdict(LOGGED_OUT_HTML, | ||
| + | if ok: | ||
| + | print(f" | ||
| + | failed += 1 | ||
| + | else: | ||
| + | print(f" | ||
| + | |||
| + | # Equality, not substring: a segment " | ||
| + | session_html = '<a href="/ | ||
| + | if " | ||
| + | print(" | ||
| + | failed += 1 | ||
| + | else: | ||
| + | print(" | ||
| + | |||
| + | aria = LOGGED_OUT_HTML.replace( | ||
| + | "</ | ||
| + | ) | ||
| + | ok, why = verdict(LOGGED_OUT_HTML, | ||
| + | if not ok: | ||
| + | print(" | ||
| + | failed += 1 | ||
| + | else: | ||
| + | print(" | ||
| + | |||
| + | if failed: | ||
| + | raise SystemExit(f" | ||
| + | |||
| + | |||
| + | def main() -> int: | ||
| + | ap = argparse.ArgumentParser(description=__doc__, | ||
| + | ap.add_argument(" | ||
| + | ap.add_argument(" | ||
| + | ap.add_argument(" | ||
| + | args = ap.parse_args() | ||
| + | |||
| + | if not args.skip_self_test: | ||
| + | self_test() | ||
| + | |||
| + | if args.logged_out is None and args.logged_in is None: | ||
| + | logged_out = LOGGED_OUT_HTML | ||
| + | logged_in = LOGGED_IN_HTML | ||
| + | print(" | ||
| + | else: | ||
| + | if args.logged_out is None or args.logged_in is None: | ||
| + | ap.error(" | ||
| + | logged_out = open(args.logged_out, | ||
| + | logged_in = open(args.logged_in, | ||
| + | print(" | ||
| + | |||
| + | ok, why = verdict(logged_out, | ||
| + | if ok: | ||
| + | print(" | ||
| + | for line in why: | ||
| + | print(" | ||
| + | return 0 | ||
| + | print(" | ||
| + | print(" | ||
| + | print(" | ||
| + | return 1 | ||
| + | |||
| + | |||
| + | if __name__ == " | ||
| + | sys.exit(main()) | ||
| </ | </ | ||
| - | === Figures === | + | Documented run, 2026-08-27, eight self-tests then the fixtures (path-segment equality, not substring; '' |
| - | To use floats, you have to use the '' | + | <code> |
| - | <WRAP right 50% box> | + | SELFTEST ok |
| - | {{PATH_TO_FILE|ALT_TEXT}} | + | logout marker present only when logged in: logout, logout-link |
| - | < | + | user marker present only when logged in: account, logged-in, user-menu |
| - | </ | + | login area present logged-out and absent logged-in: login, login-link, password |
| + | SELFTEST ok two logged-out snapshots do not verify | ||
| + | SELFTEST ok prose 'log out' | ||
| + | SELFTEST ok | ||
| + | SELFTEST ok | ||
| + | SELFTEST ok | ||
| + | SELFTEST ok ' | ||
| + | SELFTEST ok | ||
| - | */ | + | Documented run (embedded fixtures): |
| + | VERIFIED login | ||
| + | | ||
| + | user marker present only when logged in: account, logged-in, user-menu | ||
| + | login area present logged-out and absent logged-in: login, login-link, password | ||
| + | </code> | ||
| - | ====== References ====== | + | ===== SSO as a way in, versus SSO as the subject |
| - | /* | + | These are different papers. Using Facebook as a door into a ranked list is scale-login |
| - | To insert citations, follow these steps: | + | |
| - | - Verify the BibTeX entry exists | + | |
| - | - Use {[CitationKey]} where needed in the text; it will render as a numbered reference. | + | |
| - | - Keep this section unchanged to display | + | |
| + | * **SAAT**, Ghasemisharif, | ||
| - | If any step fails, | + | A crawl that only //detects// an SSO button has not logged in. Dimova et al. are explicit about that. Do not cite 7.23% as a login rate. |
| - | */ | + | |
| - | <bibtex bibliography></ | + | ===== Automated registration ===== |
| + | Five numbers people confuse. They do not share a denominator. | ||
| + | |||
| + | ^ Paper ^ What was created ^ Of what ^ Rate ^ | ||
| + | | Drakonakis et al. registered and logged in {[drakonakis2020_cookie]} | count not given | 168,594 domains with a signup option | **13.7%** | | ||
| + | | Drakonakis et al. accounts created {[drakonakis2020_cookie]} | 25,242 accounts | 168,594 signup domains | they write " | ||
| + | | Drakonakis et al., of the full crawl {[drakonakis2020_cookie]} | 25,242 accounts | 1,585,964 unique domains crawled | **~1.6%** | | ||
| + | | Al Roomi et al., USENIX Security 2023 {[alroomi2023_login]} | 45.0K initial test accounts | Google CrUX top 1M | **4.5%** | | ||
| + | | Alroomi et al., CCS 2023 {[alroomi2023_password]} | 20,119 domains fully evaluated | Tranco 1M after filtering | a completed-evaluation count, not a success rate | | ||
| + | | Kubicek et al., TheWebConf 2024 {[kubicek2024_register]} | register-or-newsletter | 660,202 unique Tranco domains | **5.9%** | | ||
| + | |||
| + | Kubicek et al. compare themselves to Cookie Hunter' | ||
| + | |||
| + | **Cookie Hunter.** XDriver on Selenium. Email verification by visiting links in mail. SMS / SSN blocked leftover of the top 1K, completed by hand; the ~25K that make the evaluation "did not require any manual intervention" | ||
| + | |||
| + | > funding human captcha-solving services to create accounts presents an ethical dilemma, we opted to not handle such cases. | ||
| + | |||
| + | **Login policies** (Al Roomi and Li) {[alroomi2023_login]}. CrUX 1M (they started ground-truth on Tranco and switched because CrUX is websites). Signup page on 258,200 domains; **45.0K** initial accounts; a second account on only 37,300. AZcaptcha. In the **ground-truth** accounts they created by hand, 39% of domains sent a verification email — that 39% is **not** of the 45.0K. They did not complete phone verification. | ||
| + | |||
| + | **Password policies** (Alroomi and Li) {[alroomi2023_password]}. Tranco 1M; **20,119** domains fully evaluated. CAPTCHAs on at least **49%** of signup forms; AZcaptcha solved **94%**. They avoided human-driven solvers "due to ethical issues identified with such services", | ||
| + | |||
| + | **Tripwire** {[deblasio2017_tripwire]}. Honey accounts to infer site compromise from password reuse. **65,413** registration attempts across **33,634** sites; **3,664** accounts on around **2,302** sites. Third-party CAPTCHA solver (DeCaptcher). "While we make no attempt to explicitly check the terms of service" | ||
| + | |||
| + | **Kubicek et al.** {[kubicek2024_register]}, | ||
| + | |||
| + | ===== Newsletter ===== | ||
| + | |||
| + | A mailing-list form is not an account. Coverage is higher, legal exposure is different, and you still have to say which one you did. | ||
| + | |||
| + | Englehardt, Han and Narayanan {[englehardt2018_email]}: | ||
| + | |||
| + | Kubíček et al. {[kubicek2022_emails]}: | ||
| + | |||
| + | > Only 59% of websites that sent us at least one email first sent us a double opt-in email. Moreover, 5.5% of services sent us an unsolicited marketing email without any confirmation or double opt-in email. | ||
| + | |||
| + | The 59% is of websites that sent mail. The 5.5% is of services that sent unsolicited marketing with no confirmation. Neither is "59% of 666". | ||
| + | |||
| + | Mathur et al. {[mathur2023_political]}: | ||
| + | |||
| + | ===== Forms without submit ===== | ||
| + | |||
| + | Senol et al. {[senol2022_leaky]} filled email on **52,055** of **99,380** loaded Tranco 100k sites (EU, no-action) and left. Schema: '' | ||
| + | |||
| + | Chatzimpyrros, | ||
| + | |||
| + | ===== CAPTCHA, email, SMS — dated ===== | ||
| + | |||
| + | CAPTCHA is **non-random loss**. Cookie Hunter' | ||
| + | |||
| + | ^ When ^ What the measurement papers actually used ^ Status in 2026 ^ | ||
| + | | 2017 (Tripwire) | third-party solver (DeCaptcher); | ||
| + | | 2020 (Cookie Hunter) | no farm; WebDriver already made Google not serve reCAPTCHA | The ethical refusal is still the argument to cite. The detection of WebDriver is not the 2026 bottleneck | | ||
| + | | 2022 crawl (Kubicek et al., published 2024) | human farm; mix **75 / 20 / 2 / 3** v2 / v3 / hCaptcha / image; later, research assistants | Dated snapshot of **what was on the web in 2022**. Not a 2026 vendor share | | ||
| + | | 2023 (Alroomi / Al Roomi) | AZcaptcha, described in those papers as OCR not humans; 94% solve rate on the password-policy crawl | **Paper-era** automated farm-API. They avoided human solvers on purpose. Live AZcaptcha docs (fetched 2026-08-27) mention a workers pool — do not treat the 2023 OCR claim as a 2026 service description | | ||
| + | | 2023 (Searles et al.) {[searles2023_captcha]} | **manual** user study, 1,400 people, 14,000 CAPTCHAs; 185 of ~200 Alexa sites had account creation, **142** succeeded | Not a crawler. Useful for "will a human complete signup" | ||
| + | | 2025 (Teoh et al., Halligan) {[teoh2025_captcha]} | agentic VLM **60.7%** of 2,600 challenges; infiltrated **2Captcha** at **70.6%** | **Current.** Farms still exist in 2025. VLMs change the " | ||
| + | | 2026 (Turnstile) | Cloudflare' | ||
| + | |||
| + | Do not take a share from Similarweb, wmtips, or any SEO page. The 75/20/2/3 mix is Kubicek et al.'s 2022 crawl. Google' | ||
| + | |||
| + | Email verification is the common path (login-policies ground-truth 39%; Cookie Hunter visited the link). SMS is the expensive path: SAAT used Twilio; Cookie Hunter treated phone/SSN as a leftover for the top 1K; login-policies did not do phone. Plan for email. Budget SMS only if the research question lives behind it. | ||
| + | |||
| + | Human CAPTCHA farms: **name who used them, and when**. Cookie Hunter no; Kubicek et al. yes then research assistants; Tripwire used DeCaptcher (a third-party solver; the paper does not establish that the solvers were human); password-policies AZcaptcha, described there as not human, on purpose; Teoh et al. infiltrated 2Captcha in 2025 as an experiment, not as a measurement instrument they recommend. Do not present a farm as current best practice. | ||
| + | |||
| + | ===== Terms of service and ethics ===== | ||
| + | |||
| + | The ethics page already ranks what crawling papers say they did. Of 697 papers that describe a mitigation, **92 (13.2%)** mention test accounts or synthetic identities, and **17 (2.4%)** mention robots.txt, terms of service or acceptable-use policies. A full-text sweep on the same 1,120 crawling papers finds 129 (11.5%) that mention terms of service at all — without separating "we complied" | ||
| + | |||
| + | What the registration papers themselves wrote: | ||
| + | |||
| + | * Tripwire {[deblasio2017_tripwire]}: | ||
| + | * Password policies {[alroomi2023_password]}: | ||
| + | * SAAT {[ghasemisharif2022_saat]}: | ||
| + | * Cookie Hunter {[drakonakis2020_cookie]}: | ||
| + | * Kubicek et al. {[kubicek2024_register]}: | ||
| + | |||
| + | There is no consensus in this literature that creating test accounts is permitted, forbidden, or ToS-exempt. The papers that created accounts at this scale wrote that they did **not** check ToS (Tripwire, password-policies) or named the tension and did it anyway (SAAT). That is five papers, not a vote of 45.0K. Write the sentence: synthetic identities, unique emails, whether you solved CAPTCHAs and how, whether you completed email/SMS, and that you did not check ToS at scale (if you did not). [[Practices: | ||
| + | |||
| + | ===== Which methods are current ===== | ||
| + | |||
| + | ^ Method ^ Status ^ | ||
| + | | Logged-out crawl of a ranked list | **Still the default, and still defensible** for questions that do not live behind an account. 553/857 are labelled '' | ||
| + | | Manual accounts, then scripted login (Kaizer 2016, To Auth 2024) | **Current** when the list is hundreds, not hundreds of thousands | | ||
| + | | Shepherd-style differential verification | **Current, and underused.** The 2020 workshop paper is out of corpus; the method is not obsolete | | ||
| + | | Cookie Hunter-style full registration on a million-site list | **Done, expensive, and CAPTCHA-censored.** The 13.7% is registered-and-logged-in of signup domains in 2020, under a no-farm constraint; the 25,242 is a separate accounts-created count | | ||
| + | | Human CAPTCHA farm as a measurement instrument | **Used (Kubicek 2022 crawl). Not best practice.** Tripwire 2015 used a third-party solver, not established as a human farm. Password-policies 2023 and Cookie Hunter 2020 both refused a human farm. Teoh 2025 shows the farm is still there to infiltrate | | ||
| + | | AZcaptcha / automated solver APIs | **Used in the 2023 policy crawls.** Date the OCR/ | ||
| + | | Agentic VLM solving (Halligan 2025) | **New.** 60.7% is a paper about whether the technique works, which is the stage it is at | | ||
| + | | Turnstile / CAPTCHA-free widgets | **Current on the web, almost unmeasured in this corpus.** Date any CAPTCHA-mix figure | | ||
| + | | Newsletter signup | **Current** as a lighter instrument; say it is not an account | | ||
| + | | Fill without submit | **Current** for exfiltration; | ||
| + | | SSO as a door into a ranked list | **Rare.** Schema SSO among web crawls is 0. SAAT and SSOScan did it; Dimova and Ghasemisharif 2018 measured SSO without completing RP login at scale | | ||
| + | |||
| + | ===== What to report ===== | ||
| + | |||
| + | A methods paragraph that answers these is enough. None of them takes more than a clause. | ||
| + | |||
| + | - **Which of the four jobs** you did (login / register / newsletter / fill-without-submit). | ||
| + | - **How many accounts, of how many sites that offered the form, of how many sites you tried.** Three numbers, not one rate. | ||
| + | - **How you verified login** — differential (Shepherd), identifier in the page, session cookie still set on a second visit. "HTTP 200 after POST" is not a verification. | ||
| + | - **CAPTCHA**: | ||
| + | - **Email and SMS**: unique addresses? visited the link? phone numbers? Twilio? | ||
| + | - **ToS / ethics**: test accounts, whether you checked ToS (usually: you did not, at this scale), IRB/legal if you have it. | ||
| + | - **SSO**: did you complete the RP login, or only count buttons? | ||
| + | |||
| + | ===== Methodology and limitations ===== | ||
| + | |||
| + | Corpus: 5,859 extracted papers, 2010–2026, | ||
| + | |||
| + | The 81 is a schema enum; the 17 is a hand map over those 81 ('' | ||
| + | |||
| + | Out of corpus, and labelled so wherever they appear: Shepherd (MADWeb 2020), Chatzimpyrros et al. (ESORICS workshop), Mathur et al. (Big Data & Society). In the bibliographic index but not extracted: Kubicek et al. WWW 2024. | ||
| + | |||
| + | The query log, the ROLE list, the rejected sweeps, and the review log are on [[provenance: | ||
| + | |||
| + | ===== Open questions ===== | ||
| + | |||
| + | <WRAP todo> | ||
| + | * Re-extract Kubicek et al. WWW 2024 so its figures sit in '' | ||
| + | * A 2026 CAPTCHA-provider mix on CrUX or Tranco, with Turnstile as a first-class label. The 75/20/2/3 snapshot is 2022. | ||
| + | * How often a Shepherd-style verifier disagrees with "the login form disappeared" | ||
| + | * A hand split of the 129 crawling papers that mention terms of service into "we complied" | ||
| + | </ | ||
| + | |||
| + | ===== Related pages ===== | ||
| + | |||
| + | * [[Programming: | ||
| + | * [[Programming: | ||
| + | * [[Privacy: | ||
| + | * [[Practices: | ||
| + | * [[Programming: | ||
| + | * [[Programming: | ||
| + | |||
| + | ====== References ====== | ||
| + | |||
| + | <bibtex bibliography></ | ||
| - | /* This enables discussion under this article. */ | ||
| ~~DISCUSSION~~ | ~~DISCUSSION~~ | ||
| + | |||
programming/registration.1739892438.txt.gz · Last modified: by karelkubicek
