| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| design:platforms [2026/08/27 20:25] – Note that the platform-family rows do not sum to 897 (a paper can name several); escape two published regexes as ''%%..%%''. Authored by Claude karel.kubicek.claude | design:platforms [2026/08/27 20:51] (current) – Quote the papers' own contiguous sentences for the Mastodon 96% and the WhatsApp candidate-number count (both were paraphrases of column-spliced text). Authored by Claude karel.kubicek.claude |
|---|
| |
| <WRAP important> | <WRAP important> |
| **The access route is the design decision, and the routes changed under the field.** In the [[literature:corpus|publication corpus]] (seven venues, 2010–2026, **5,859** extracted papers), **897** papers name a large platform as the subject of measurement — **15.7%** of the **5,712** that drew any study population. Of those 897, **383 (42.7%)** worked from an **existing dataset** rather than collecting anything themselves — a higher share than live collection (**346, 38.6%**). | **The access route is the design decision, and the routes changed under the field.** In the [[literature:corpus|publication corpus]] (seven venues, 2010–2026, **5,859** extracted papers), **897** papers name a large platform as the subject of measurement — **15.7%** of the **5,712** that drew any study population. Of those 897, **383 (42.7%)** name an **existing dataset** as a source of their data, and **208 (23.2%)** name one and no primary collection of their own at all. |
| |
| Meanwhile the official researcher routes that replaced the open APIs are **almost entirely absent from this literature**. Across the **5,855** papers with full text: the **Meta Content Library** appears **once**, and that once is a parenthetical noting that CrowdTangle "//now replaced by the Meta Content Library//" {[roy2025_darkgram]}. The **TikTok Research API** appears in **zero** papers. The **YouTube Researcher Program** appears in **two**. The X/Twitter **Academic Research track** appears in **eight**. By contrast **Pushshift** — a third-party Reddit archive — appears in **35**, and **Ad Library** in **69**. | Meanwhile the official researcher routes that replaced the open APIs are **almost entirely absent from this literature**. Across the **5,855** papers with full text: the **Meta Content Library** appears **once**, and that once is a parenthetical noting that CrowdTangle "//now replaced by the Meta Content Library//" {[roy2025_darkgram]}. The **TikTok Research API** appears in **zero** papers. The **YouTube Researcher Program** appears in **two**. The X/Twitter **Academic Research track** appears in **eight**. By contrast **Pushshift** — a third-party Reddit archive — appears in **35**, and **Ad Library** in **69**. |
| |
| If you plan a platform study in 2026 through an official research programme, assume you are an early adopter: there is no methods section in these venues to copy, and the application lead time is weeks, not days. | Part of that absence is **mechanical**, and the page says so rather than letting you infer otherwise: these programmes are 2023-and-later launches read through a corpus whose 2025–2026 slice is under-covered by construction, and the venues that publish most platform work of this kind — ICWSM, CHI, FAccT, the communications journals — are **not in the corpus at all**. The right reading is "there is no worked example //in these seven venues//", not "the programmes do not work". |
| | |
| | Either way the consequence for you is the same. If you plan a platform study in 2026 through an official research programme, assume you are an early adopter: there is no methods section in these venues to copy, and the application lead time is weeks, not days. |
| </WRAP> | </WRAP> |
| |
| * {[vombatkere2024_tiktok]} (TheWebConf 2024) — **sock-puppet accounts** on a feed you cannot query. Five bot accounts plus a donated dataset of 347 real users; concludes that TikTok "//exploits real users' interests in between 30% and 50% of all recommended videos in the first thousand videos//". Read it for how you measure a recommender with no API at all. | * {[vombatkere2024_tiktok]} (TheWebConf 2024) — **sock-puppet accounts** on a feed you cannot query. Five bot accounts plus a donated dataset of 347 real users; concludes that TikTok "//exploits real users' interests in between 30% and 50% of all recommended videos in the first thousand videos//". Read it for how you measure a recommender with no API at all. |
| * {[karnam2026_setting]} (IEEE S&P 2026) — the **data subject's right of access** as the instrument. Sock-puppet accounts on Instagram, TikTok and YouTube, then GDPR Art. 15 requests, then a comparison of the returned data package against the ground truth the authors themselves generated. All three platforms failed to report purpose, recipients and retention period. | * {[karnam2026_setting]} (IEEE S&P 2026) — the **data subject's right of access** as the instrument. Sock-puppet accounts on Instagram, TikTok and YouTube, then GDPR Art. 15 requests, then a comparison of the returned data package against the ground truth the authors themselves generated. All three platforms failed to report purpose, recipients and retention period. |
| * {[gegenhuber2026_there]} (NDSS 2026) — **enumeration at platform scale**, and the ethics that go with it. Probes 63 billion candidate phone numbers and discovers **3,546,479,731** WhatsApp accounts across 245 countries, **57%** of them with a public profile picture. Read it for what a platform-scale study has to say about disclosure and harm before a programme committee will accept it. | * {[gegenhuber2026_there]} (NDSS 2026) — **enumeration at platform scale**, and the ethics that go with it. Probes 63,170,000,000 candidate phone numbers and discovers **3,546,479,731** WhatsApp accounts across 245 countries, **57%** of them with a public profile picture. Read it for what a platform-scale study has to say about disclosure and harm before a programme committee will accept it. |
| |
| If you only read one thing about the **rules**, read the Commission's December 2025 decision fining X €120 million: it is the first DSA non-compliance decision, and one of the three grounds is research data access.((European Commission, //Commission fines X €120 million under the Digital Services Act//, press release, 5 December 2025, https://digital-strategy.ec.europa.eu/en/news/commission-fines-x-eu120-million-under-digital-services-act — page fetched 2026-08-27, "Last update 16 January 2026".)) | If you only read one thing about the **rules**, read the Commission's December 2025 decision fining X €120 million: it is the first DSA non-compliance decision, and one of the three grounds is research data access.((European Commission, //Commission fines X €120 million under the Digital Services Act//, press release, 5 December 2025, https://digital-strategy.ec.europa.eu/en/news/commission-fines-x-eu120-million-under-digital-services-act — page fetched 2026-08-27, "Last update 16 January 2026".)) |
| Every route below is real and in use somewhere. What differs is whether it is current, what it costs, and what population it yields. Dates are as of **2026-08-27** and verified against primary sources; the [[provenance:design:platforms|provenance page]] lists each URL, quote and fetch date, and names the claims that could not be verified. | Every route below is real and in use somewhere. What differs is whether it is current, what it costs, and what population it yields. Dates are as of **2026-08-27** and verified against primary sources; the [[provenance:design:platforms|provenance page]] lists each URL, quote and fetch date, and names the claims that could not be verified. |
| |
| ^ Route ^ Status in 2026 ^ What it gets you ^ Papers in our corpus that name it ^ | The **Status** column is one word; the dated evidence for each verdict is in [[#Which Methods Are Current]] at the foot of the page. |
| | **Open public API, free** | **historical** for the platforms this page is about | the population the platform chose to expose | 91 papers name a Twitter/X streaming or search API, peaking 2022 | | |
| | **Paid metered API** | **current** — X moved to pay-per-usage | whatever you can afford | see the arithmetic below | | ^ Route ^ Status ^ What it gets you ^ Papers in our corpus that name it ^ |
| | **Platform research programme** | **current, and unused in this literature** | a curated, filtered archive inside a controlled environment | Meta Content Library **1**, TikTok Research API **0**, YouTube Researcher Program **2** | | | **Open public API, free** | **historical** | the population the platform chose to expose | 91 papers name a Twitter/X streaming or search API, peaking 2022 | |
| | **Regulator route (DSA Art. 40 vetted researcher)** | **current, newly operational** | in principle non-public data from a designated VLOP | effectively **zero** papers; see the audit below | | | **Paid metered API** | **current** | whatever you can afford | X moved to pay-per-usage; see the arithmetic below | |
| | | **Platform research programme** | **current** | a curated, filtered archive inside a controlled environment | Meta Content Library **1**, TikTok Research API **0**, YouTube Researcher Program **2** | |
| | | **Regulator route (DSA Art. 40 vetted researcher)** | **current** | in principle non-public data from a designated VLOP | effectively **zero** papers; see the audit below | |
| | **Transparency / ad archive** | **current** | only what the archive covers — usually ads that actually ran | **69** papers name an Ad Library or ad archive | | | **Transparency / ad archive** | **current** | only what the archive covers — usually ads that actually ran | **69** papers name an Ad Library or ad archive | |
| | **Third-party archive of platform data** | **current in practice** | someone else's crawl, with their gaps | **35** papers name Pushshift | | | **Third-party archive of platform data** | **historical** | someone else's crawl, with their gaps — the dumps circulate, the live service does not | **35** papers name Pushshift | |
| | **Reuse of a published platform dataset** | **current, and dominant** | a frozen snapshot you did not design | **383 of 897** platform papers are `existing-dataset` | | | **Reuse of a published platform dataset** | **current** | a frozen snapshot you did not design | **383 of 897** name an existing dataset; **208** name one //and no primary collection// | |
| | **Logged-out scraping** | **current, legally contested — and now the subject of a DSA fine against a platform that forbade it** | the public surface, as a bot sees it | **157** papers name a scraping or automation tool | | | **Logged-out scraping** | **current, contested** | the public surface, as a bot sees it | **157** papers name a scraping or automation tool | |
| | **Accounts you created (sock puppets)** | **current, highest-risk** | the logged-in surface, and anything personalised | **14** papers use the term "sock puppet"; **22** state `account-registration` | | | **Accounts you created (sock puppets)** | **current** | the logged-in surface, and anything personalised | **14** papers use the term "sock puppet"; **22** are labelled `account-registration` | |
| | **Data donation** | **current and rising** | real users' own data, with consent | **25** papers mention data donation, **11** of them in 2025 | | | **Data donation** | **current** | real users' own data, with consent | **25** papers mention data donation, **11** of them in 2025 | |
| | **Right of access / DSAR** | **current and rising** | what the platform says it holds about a subject | **52** papers, rising from 2022 | | | **Right of access / DSAR** | **current** | what the platform says it holds about a subject | **52** papers, rising from 2022 | |
| | **On-device or client-side extraction** | **niche, powerful** | the model or logic actually shipped to users | {[west2024_picture]} extracted the on-device ML models from the Instagram and TikTok apps | | | **On-device or client-side extraction** | **niche** | the model or logic actually shipped to users | {[west2024_picture]} extracted the on-device ML models from the Instagram and TikTok apps | |
| | **Operator collaboration** | **current, not replicable** | everything, and no reproducibility | {[cohn2020_delf]}, {[schlinker2019_internet]} | | | **Operator collaboration** | **not replicable** | everything, and no reproducibility | {[cohn2020_delf]}, {[schlinker2019_internet]} | |
| |
| ==== The metered API changes the study design, not just the budget ==== | ==== The metered API changes the study design, not just the budget ==== |
| |
| X's API is now **pay-per-usage with no subscription**: %%$0.005%% per Post read, %%$0.010%% per User read, and "//Pay-per-usage plans are capped at 3 million Post reads per monthly billing cycle//" with Enterprise above that.((https://docs.x.com/x-api/getting-started/pricing — fetched 2026-08-27; the page carries no date. There is no free tier and no academic tier in the current documentation tree.)) One million posts is therefore **%%$5,000%%** and the monthly ceiling on the self-serve route is **3 million posts**. A study that used to be "collect the 1% stream for six months" is now a budget line and a cap. | X's API is now **pay-per-usage with no subscription**: %%$0.005%% per Post read, %%$0.010%% per User read, and "//Pay-per-usage plans are capped at 3 million Post reads per monthly billing cycle//" with Enterprise above that.((https://docs.x.com/x-api/getting-started/pricing — fetched 2026-08-27; the page carries no date. On the same fetch, the //About the X API// page lists only v2 (pay-per-usage) and v1.1 ("limited support"), and neither it nor the pricing page mentions a free or academic tier; we did not exhaustively search the whole documentation tree, so read that as "not offered on the pages a new applicant is sent to" rather than as a proof of absence.)) One million posts is therefore **%%$5,000%%** and the monthly ceiling on the self-serve route is **3 million posts**. A study that used to be "collect the 1% stream for six months" is now a budget line and a cap. |
| |
| This is not hypothetical, and the field has already adapted. {[nguyen2025_please]} (USENIX Security 2025) built its collection pipeline around the read cap: it polls a **counts** endpoint, which "//does not add to our monthly tweet read limit//", and only fetches actual posts when the count is non-zero — because "//the total number of tweets that can be downloaded is very small for the Basic tier//". Note that even //that// adaptation is already dated: the Basic tier the paper worked around no longer exists. | This is not hypothetical, and the field has already adapted. {[nguyen2025_please]} (USENIX Security 2025) built its collection pipeline around the read cap: it polls a **counts** endpoint, which "//does not add to our monthly tweet read limit//", and only fetches actual posts when the count is non-zero — because "//the total number of tweets that can be downloaded is very small for the Basic tier//". That is still a sound pattern, but the tier it names is no longer what a new project gets: X's own changelog records that pay-per-usage "//officially launched//" on **6 February 2026** as the self-serve model, with the older Basic and Pro plans "//remain[ing] available//" only to existing subscribers, and then a further price change on **20 April 2026** introducing "//Owned Reads//" at %%$0.001%% per resource for a developer's own account data.((https://docs.x.com/changelog — fetched 2026-08-27; entries dated 6 February 2026 ("Launch of X API Pay-Per-Use pricing") and 16 April 2026 ("X API pricing update: Owned Reads now %%$0.001%%").)) The pricing model of the single most-measured platform in this corpus changed **twice inside 2026**. Date it, and re-check it before you submit. |
| |
| The other side of the same story: {[galeazzi2026_revealing]} (NDSS 2026) states that academic access to the X API has "//been restricted since June 2023//" and therefore works from two **previously published** tweet datasets rather than collecting anything. Both papers are correct methodology for their moment. Neither is a template you can lift. | The other side of the same story: {[galeazzi2026_revealing]} (NDSS 2026) states that academic access to the X API has "//been restricted since June 2023//" and therefore works from two **previously published** tweet datasets rather than collecting anything. Both papers are correct methodology for their moment. Neither is a template you can lift. |
| |
| * **Meta Content Library and Content Library API** — covers public content from Facebook, Instagram, WhatsApp Channels and (in the web tool) Threads. Applications are reviewed **not by Meta** but independently by the **Secure Data Access Center (CASD)** in France. Eligibility is an accredited, degree-granting, not-for-profit academic institution, or a not-for-profit research institution.((https://transparency.meta.com/researchtools/meta-content-library/ ("UPDATED APR 30, 2026") and https://developers.facebook.com/docs/content-library-and-api/get-access — both fetched 2026-08-27.)) | * **Meta Content Library and Content Library API** — covers public content from Facebook, Instagram, WhatsApp Channels and (in the web tool) Threads. Applications are reviewed **not by Meta** but independently by the **Secure Data Access Center (CASD)** in France. Eligibility is an accredited, degree-granting, not-for-profit academic institution, or a not-for-profit research institution.((https://transparency.meta.com/researchtools/meta-content-library/ ("UPDATED APR 30, 2026") and https://developers.facebook.com/docs/content-library-and-api/get-access — both fetched 2026-08-27.)) |
| * **TikTok Research Tools** — accounts, content and Shops data; open to academic institutions in the US, EEA, UK or Switzerland, to EU not-for-profit or independent research bodies, and (for youth-safety research only) to Brazil. Stated turnaround "//within 4 weeks//". The separate **Commercial Content Library** covers advertising and, as of this fetch, "//we are ONLY including data from EU countries//".((https://developers.tiktok.com/products/research-api/ and https://developers.tiktok.com/products/commercial-content-api/ — fetched 2026-08-27.)) | * **TikTok Research Tools** — accounts, content and Shops data; open to "//Academic institutions in the US, EEA, UK or Switzerland//", to a "//Not-for-profit and/or independent research institution, organization, association, or body in the EU//" — immediately followed on the same page by "//We are currently beta testing this service with select researchers in the US, UK, Switzerland, Norway, Iceland and Liechtenstein//", so read the non-academic pathway as a beta rather than as open — and (for youth-safety research only) to Brazil. Stated turnaround "//within 4 weeks//". The separate **Commercial Content Library** covers advertising and, as of this fetch, "//we are ONLY including data from EU countries//".((https://developers.tiktok.com/products/research-api/ and https://developers.tiktok.com/products/commercial-content-api/ — fetched 2026-08-27.)) |
| * **YouTube Researcher Program** — "//scaled, expanded access to global video metadata across the entire public YouTube corpus via our Data API//", for students, research staff and faculty at accredited institutions. Separate from, and not satisfied by, the ordinary Data API quota-increase form.((https://research.youtube/how-it-works/ — fetched 2026-08-27. The default Data API v3 allocation is 10,000 units/day plus 100 `search.list` calls: https://developers.google.com/youtube/v3/getting-started#quota.)) | * **YouTube Researcher Program** — "//scaled, expanded access to global video metadata across the entire public YouTube corpus via our Data API//", for students, research staff and faculty at accredited institutions. Separate from, and not satisfied by, the ordinary Data API quota-increase form.((https://research.youtube/how-it-works/ — fetched 2026-08-27. The default Data API v3 allocation is 10,000 units/day plus 100 `search.list` calls: https://developers.google.com/youtube/v3/getting-started#quota.)) |
| |
| ==== The regulator route: new, and untested in this literature ==== | ==== The regulator route: new, and untested in this literature ==== |
| |
| For services the European Commission has designated as **very large online platforms** (VLOPs), Article 40 of the Digital Services Act creates a **vetted-researcher** data-access right. The implementing instrument is **Commission Delegated Regulation (EU) 2025/2050 of 1 July 2025**, which lays down "//the technical conditions and procedures under which providers of very large online platforms and of very large online search engines are to share data with vetted researchers//" and enters into force on the twentieth day after its publication in the Official Journal.((https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=OJ:L_202502050 (CELEX 32025R2050) — fetched 2026-08-27. Vetting is done by a national Digital Services Coordinator; the Commission runs a portal at https://data-access.dsa.ec.europa.eu/ (JavaScript-only, title "DSA - Data Access Portal", fetched 2026-08-27).)) The current designation list is maintained by the Commission and moves: it was last updated **24 July 2026** and now includes WhatsApp Ireland alongside Facebook, Instagram, TikTok, YouTube, X and the Amazon Store.((https://digital-strategy.ec.europa.eu/en/policies/list-designated-vlops-and-vloses — fetched 2026-08-27, "Information updated on 24 July 2026".)) | For services the European Commission has designated as **very large online platforms** (VLOPs), Article 40 of the Digital Services Act creates a **vetted-researcher** data-access right. The implementing instrument is **Commission Delegated Regulation (EU) 2025/2050 of 1 July 2025**, which lays down "//the technical conditions and procedures under which providers of very large online platforms and of very large online search engines are to share data with vetted researchers//" and enters into force on the twentieth day after its publication in the Official Journal — which was **9 October 2025**, so it has been in force since **29 October 2025**.((https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=OJ:L_202502050 (CELEX 32025R2050) — fetched 2026-08-27; the record reads "OJ L, 2025/2050, 9.10.2025".)) Vetting is done by a national Digital Services Coordinator, and applications go through the Commission's **DSA Data Access Portal**, which states "//You can send applications as of 29 October 2025//" and carries a public **Research projects** register of "//ongoing research projects conducted by vetted researchers who have access to data under Article 40//".((https://data-access.dsa.ec.europa.eu/home — JavaScript-only; fetched 2026-08-27 with a headless browser. That register, not this corpus, is where the first worked example will appear.)) The current designation list is maintained by the Commission and moves: it was last updated **24 July 2026** and now includes WhatsApp Ireland alongside Facebook, Instagram, TikTok, YouTube, X and the Amazon Store.((https://digital-strategy.ec.europa.eu/en/policies/list-designated-vlops-and-vloses — fetched 2026-08-27, "Information updated on 24 July 2026".)) |
| |
| <WRAP tip> | <WRAP tip> |
| * **A research archive is not the platform.** The Meta Content Library includes Facebook posts to Pages, groups and events "//as well as posts that appear on public profiles that are either verified or that have 100 or more followers//"; Instagram business and creator accounts plus personal accounts "//verified or that have 100 or more followers//"; Threads public profiles with 100 or more followers.((https://transparency.meta.com/researchtools/meta-content-library/, fetched 2026-08-27.)) A prevalence computed there is a prevalence **among accounts above a follower threshold**, not among users. | * **A research archive is not the platform.** The Meta Content Library includes Facebook posts to Pages, groups and events "//as well as posts that appear on public profiles that are either verified or that have 100 or more followers//"; Instagram business and creator accounts plus personal accounts "//verified or that have 100 or more followers//"; Threads public profiles with 100 or more followers.((https://transparency.meta.com/researchtools/meta-content-library/, fetched 2026-08-27.)) A prevalence computed there is a prevalence **among accounts above a follower threshold**, not among users. |
| * **An ad archive contains ads that ran.** It cannot tell you what was rejected, and it can be wrong in both directions about what is political — which is exactly {[bouchaud2024_beyond]}'s finding, on both sides at once. | * **An ad archive contains ads that ran.** It cannot tell you what was rejected, and it can be wrong in both directions about what is political — which is exactly {[bouchaud2024_beyond]}'s finding, on both sides at once. |
| * **A sampled stream is a sample.** {[hagen2021_numbers]} is explicit that its dataset came from "//the 1% streaming API that Twitter provides to vetted researchers//", so all of its figures are lower bounds. Say which stream and which sampling rate. | * **A sampled stream is a sample.** {[kaleli2021_human]} is explicit that its dataset came from "//the 1% streaming API that Twitter provides to vetted researchers//" and that consequently "//all the numbers that we presented in this paper are lower bounds//". Say which stream and which sampling rate, and say which direction the bias runs. |
| * **A third-party archive has the archiver's gaps, not yours.** Pushshift is still the most-named Reddit source in this corpus (**35** papers, e.g. {[zannettou2018_origins]}), and its coverage window and completeness are properties of Pushshift, not of Reddit. | * **We could not establish what Reddit currently offers a researcher.** Every Reddit-owned domain refused our fetches, so this page makes **no claim** about Reddit's API pricing, rate limits or researcher terms — only about what the corpus shows (63 papers measuring Reddit, 14 naming a Reddit API or PRAW, 35 naming Pushshift). Check it yourself before planning around it, and see [[provenance:design:platforms]] for exactly what failed. |
| * **A reused dataset freezes a platform that has since changed.** **383 of the 897** platform-subject papers are `existing-dataset`. If that is you, date the snapshot and say what changed on the platform since — the four false positives in our Twitter/X audit were all papers using Twitter-derived benchmark corpora with no relationship to the platform at collection time. | * **A third-party archive has the archiver's gaps, not yours.** Pushshift is still the most-named Reddit source in this corpus (**35** papers, e.g. {[zannettou2018_origins]}), and its coverage window and completeness are properties of Pushshift, not of Reddit. It is also **no longer a route you can open**: Reddit's own moderator help page says Pushshift API access "//will be reinstated for verified Reddit moderators//", that each one "//need[s] explicit approval from Reddit//", and that "//the use of Pushshift will be limited to moderation use cases only//".((https://support.reddithelp.com/hc/en-us/articles/16470271632404-Pushshift-Access-Request — fetched 2026-08-27 with a headless browser (the page is Cloudflare-walled to ''%%curl%%''); the page is stamped "Updated 1 year ago". A paper citing Pushshift today is citing a historical dump, and should say which one and when it was obtained.)) |
| * **A cross-platform study has one denominator per platform, and they are not comparable.** {[acharya2024_imitation]} looked for brand-impersonation accounts targeting the top 10K Tranco brands on X, Instagram, Telegram and YouTube at once; {[beluri2025_exploration]} tracked accounts advertised for sale across five platforms and found blocking efficacy ranging from **5.02%** (YouTube) to **46.41%** (Instagram). A cross-platform rate is only meaningful if you say what you could see on each platform, because the routes in differ and so does the visible surface. A gap between two platforms may be a gap between two access routes. | * **A reused dataset freezes a platform that has since changed.** **383 of the 897** platform-subject papers name an existing dataset (''temporal.mode'' is multi-valued, so 175 of those also collected something themselves; **208** collected nothing). If that is you, date the snapshot and say what changed on the platform since — the four false positives in our Twitter/X audit were all papers using Twitter-derived benchmark corpora with no relationship to the platform at collection time. |
| * **The logged-out surface is not the platform either.** Of the **260** platform-subject papers that report a crawl configuration, **145 (55.8%)** state that they used **no authentication at all**, only **22** registered an account, **4** logged in manually, and **none** used automated login or SSO. Whatever those studies measured, it is what an anonymous visitor sees. | * **A cross-platform study has one denominator per platform, and they are not comparable.** {[acharya2024_imitation]} looked for brand-impersonation accounts targeting the top 10K Tranco brands on X, Instagram, Telegram and YouTube at once; {[beluri2025_exploration]} tracked accounts advertised for sale across five platforms and found blocking efficacy ranging from **5.02%** (YouTube) up to **48%** — its own summary is that "//TikTok and Instagram demonstrated the highest detection efficacy at 48%, whereas YouTube and Facebook showed the lowest efficacy at just 5%//". A cross-platform rate is only meaningful if you say what you could see on each platform, because the routes in differ and so does the visible surface. A gap between two platforms may be a gap between two access routes. |
| | * **The logged-out surface is not the platform either.** Of the **260** platform-subject papers that report a crawl configuration, **145 (55.8%)** are labelled as using **no authentication at all**, **22** as registering an account, **4** as logging in manually, and **none** as using automated login or SSO. Whatever those studies measured, it is mostly what an anonymous visitor sees. **These are schema labels, not audited ones**, and [[Programming:Registration]] is explicit about why that matters for this exact field: ''%%crawlConfig%%'' carries **one evidence quote for the whole object**, so the ''%%authentication%%'' label cannot be checked against its quote, and a hand-audit of **another field of that same object**, ''%%consentAction%%'', found 19.4% false positives ([[Privacy:consent]]). We did not audit these 145. Treat the shape as real and the precise share as unverified. |
| |
| ===== Rate Limits, Quotas, Bans and the Login Wall ===== | ===== Rate Limits, Quotas, Bans and the Login Wall ===== |
| | sock puppet | 14 (**1.6%**) | 26 (0.4%) | 3 of 108 (2.8%) | | | sock puppet | 14 (**1.6%**) | 26 (0.4%) | 3 of 108 (2.8%) | |
| |
| Read the ToS row as the headline: a platform paper is roughly **twice as likely** as an average paper in these venues to discuss terms of service, and in 2025 more than one in five did. That is the direction the reviewing is moving. | Read the ToS row as the headline: at **14.9%**, a platform paper is roughly **twice as likely as the average paper in these seven venues** (7.8%) to discuss terms of service, and in 2025 more than one in five did. Against a narrower and fairer baseline — papers that ran a crawl, where [[Programming:Registration]] reports 11.5% — the ratio is about 1.3×. Either way the direction is the same, and it is the direction the reviewing is moving. |
| |
| Three practical consequences: | Four practical consequences: |
| |
| * **Design for the limit, not around it.** {[nguyen2025_please]}'s counts-endpoint trigger is the pattern: find the cheap query that tells you whether the expensive query is worth making. Report the limit you worked under, because it bounds your recall. | * **Design for the limit, not around it.** {[nguyen2025_please]}'s counts-endpoint trigger is the pattern: find the cheap query that tells you whether the expensive query is worth making. Report the limit you worked under, because it bounds your recall. |
| * **The limit can be the finding.** {[xue2021_throttling]} measured Russian ISPs throttling Twitter and showed the mechanism was SNI-based: "//throttling is triggered upon observing Twitter-related domains (*.twimg.com, twitter.com, t.co) in the SNI//". If you are being throttled and can characterise by what, that is a result. | * **The limit can be the finding.** {[xue2021_throttling]} measured Russian ISPs throttling Twitter and showed the mechanism was SNI-based: "//throttling is triggered upon observing Twitter-related domains (*.twimg.com, twitter.com, t.co) in the SNI//". If you are being throttled and can characterise by what, that is a result. |
| * **Expect the platform to fight your crawler, and say whether it did.** Only **14 of 260** platform-subject papers with a crawl configuration (**5.4%**) say anything about robots.txt — barely above the **4.7%** baseline for all papers with a crawl configuration. If the platform served you a bot challenge, that is a measurement result about the platform, not an embarrassment. | * **Expect the platform to fight your crawler, and say whether it did.** Only **14 of 260** platform-subject papers with a crawl configuration (**5.4%**) say anything about robots.txt — barely above the **4.7%** baseline for all papers with a crawl configuration. If the platform served you a bot challenge, that is a measurement result about the platform, not an embarrassment. |
| * **Sock puppets are the highest-risk route.** They get you the personalised surface — {[vombatkere2024_tiktok]} and {[karnam2026_setting]} both need them — and they are the fact pattern most likely to breach terms of service and, in the US, to have been litigated. Get ethics review first (see [[Practices:Ethics]]), keep the accounts minimal, and write down what you did: **only 32.5%** of the **841** empirical platform-subject papers state an ethics-review outcome, statistically indistinguishable from the **33.8%** corpus baseline. That is not good enough for a study whose method is creating fake accounts on someone else's service. | * **Sock puppets are the highest-risk route.** They get you the personalised surface — {[vombatkere2024_tiktok]} and {[karnam2026_setting]} both need them — and they are the fact pattern most likely to breach terms of service and, in the US, to have been litigated. Get ethics review first (see [[Practices:Ethics]]), keep the accounts minimal, and write down what you did: **only 32.5%** of the **841** empirical platform-subject papers state an ethics-review outcome — essentially at the **33.8%** corpus baseline (no test was run, and the platform papers are inside that baseline). That is not good enough for a study whose method is creating fake accounts on someone else's service. |
| |
| ==== Terms of service are no longer only the platform's weapon ==== | ==== Terms of service are no longer only the platform's weapon ==== |
| | TikTok | 7 | 7 (all) | 4 — {[west2024_picture]}, {[vombatkere2024_tiktok]}, {[simko2024_modern]}, {[karnam2026_setting]} | **57%** | two interview studies where TikTok is the topic, one tool-list collision | | | TikTok | 7 | 7 (all) | 4 — {[west2024_picture]}, {[vombatkere2024_tiktok]}, {[simko2024_modern]}, {[karnam2026_setting]} | **57%** | two interview studies where TikTok is the topic, one tool-list collision | |
| | Twitter/X | 142 | 15 (every 10th) | 11 | **73%** | all four were reused Twitter-derived **benchmark corpora** (Twitter15/16, CrisisMMD), and all four were 2022 or later | | | Twitter/X | 142 | 15 (every 10th) | 11 | **73%** | all four were reused Twitter-derived **benchmark corpora** (Twitter15/16, CrisisMMD), and all four were 2022 or later | |
| | Meta | 151 | 15 (every 10th) | 10 | **67%** | Facebook Hateful Memes benchmark, Facebook-as-an-identity-provider example, related-work mention | | | Meta | 151 | 16 (every 10th) | 11 | **69%** | Facebook Hateful Memes benchmark, Facebook-as-an-identity-provider example, related-work mention | |
| | Amazon | 72 | 11 (every 7th) | 4 | **36%** | Amazon Reviews benchmarks, devices //bought on// Amazon, a hosted speech service | | | Amazon | 72 | 11 (every 7th) | 4 | **36%** | Amazon Reviews benchmarks, devices //bought on// Amazon, a hosted speech service | |
| |
| Treat the ranking below as a ranking. Do not quote a family's count as a precise number of papers without the precision above attached to it. | Treat the ranking below as a ranking. Do not quote a family's count as a precise number of papers without the precision above attached to it — and note that these samples are **small**: at //n// = 15 the 95% interval around 73% is roughly ±20 points. They are coarse corrections, not measurements. |
| |
| ==== Which platforms the field measures ==== | ==== Which platforms the field measures ==== |
| | 19 | eBay | 6 | 0.1% | | | 19 | eBay | 6 | 0.1% | |
| |
| The top two rows are largely **app-store** work, which belongs to [[Design:Mobile and app measurement]] rather than here. Note the last rows: **TikTok is 7 papers**, of which 4 survived a full audit. There is no body of TikTok measurement in these seven venues to systematise. | Rows 1 and 7 — Google and the Apple App Store — are largely **app-store** work, which is closer to [[Design:Mobile and app measurement]] than to this page. Note also what this page does **not** cover: search-engine result auditing and the Google ads ecosystem are a real body of work inside these venues and have no page on this wiki yet; they are out of scope here and named in [[#Open Questions]]. Note the last rows: **TikTok is 7 papers**, of which 4 survived a full audit. There is no body of TikTok measurement in these seven venues to systematise. |
| |
| ==== Amazon means five different things, and mostly not the shop ==== | ==== Amazon means five different things, and mostly not the shop ==== |
| |
| A substring match on "Amazon" therefore over-counts Amazon platform measurement by roughly **sixteen to one**. An earlier iteration of our own script did exactly this and reported 176 Amazon papers. | A substring match on "Amazon" therefore over-counts Amazon platform measurement by roughly **sixteen to one**. An earlier iteration of our own script did exactly this and reported 176 Amazon papers. |
| | |
| | <WRAP tip> |
| | **Two other pages count these same names and get different numbers. Both are right; they answer different questions.** [[Design:Website selection]] reports **463** papers for Alexa, folded over the 1,143 that drew a **web-unit** population. The **412** above is the subset our role rule assigns to the ranking list, over all 5,492 papers with a stated population source. Reconciling them: **486** papers name "Alexa" in a stated population source at all, **464** of those also have a web-unit population tuple — which is the figure [[Design:Website selection]] is reporting — and 412 is what survives after Alexa-as-skill-store strings are diverted to the ''%%subject%%'' role. Likewise our **179** Mechanical Turk papers is a count of ''%%population.sourceList%%'' strings; [[Design:User studies]] publishes three //different// bounds on the same thing — 139 by the schema enum, 92 by quote, 279 by full text — and warns against reading any of them as usage. Neither fold has been reconciled with the other, and this page does not claim its number is the better one. |
| | </WRAP> |
| |
| ==== The Twitter/X curve, and what it does and does not show ==== | ==== The Twitter/X curve, and what it does and does not show ==== |
| | Meta / Facebook Ad Library | 5 | 0.6% | | | Meta / Facebook Ad Library | 5 | 0.6% | |
| | CrowdTangle | 5 | 0.6% | | | CrowdTangle | 5 | 0.6% | |
| | TikTok API / TikTok-Api | 1 | 0.1% | | | ''%%TikTok-Api%%'' — the **unofficial** scraper library, not the Research API | 1 | 0.1% | |
| | **named no tool for it at all** | **552** | **61.5%** | | | **named no tool for it at all** | **552** | **61.5%** | |
| |
| That last row is the reporting gap on this page: **61.5%** of platform-subject papers name no instrument by which they reached the platform. The route is the design decision, and three papers in five do not state it. | That last row is the reporting gap on this page: for **61.5%** of platform-subject papers, **no named instrument survives into this corpus's tool lists**. Read that as an upper bound on the gap rather than as "three in five do not state it": extraction recall on free-text tool names is imperfect, and we did not hand-audit the 552. Even as an upper bound it is the largest single reporting hole the page found, because the route is the design decision. |
| |
| ==== The reproducibility cost ==== | ==== The reproducibility cost ==== |
| ^ Finding ^ Denominator the paper used ^ | ^ Finding ^ Denominator the paper used ^ |
| | Only **7.7%** of undeclared political ads were moderated as political by Meta; **60.4%** of the ads Meta did moderate did not match its own criteria {[bouchaud2024_beyond]} | 29.5 M ads from the Meta Ad Library API, 16 EU countries | | | Only **7.7%** of undeclared political ads were moderated as political by Meta; **60.4%** of the ads Meta did moderate did not match its own criteria {[bouchaud2024_beyond]} | 29.5 M ads from the Meta Ad Library API, 16 EU countries | |
| | **3,546,479,731** WhatsApp accounts discovered; **57%** with a public profile picture, **66%** of a 500,000-image sample containing a detectable face {[gegenhuber2026_there]} | 63.2 bn candidate phone numbers enumerated | | | **3,546,479,731** WhatsApp accounts discovered; **57%** with a public profile picture, **66%** of a 500,000-image sample containing a detectable face {[gegenhuber2026_there]} | 63,170,000,000 candidate phone numbers enumerated | |
| | TikTok "//exploits real users' interests in between 30% and 50% of all recommended videos in the first thousand videos//" {[vombatkere2024_tiktok]} | 347 donating users, 4.9 M videos, plus 5 bot accounts | | | TikTok "//exploits real users' interests in between 30% and 50% of all recommended videos in the first thousand videos//" {[vombatkere2024_tiktok]} | 347 donating users, 4.9 M videos, plus 5 bot accounts | |
| | **17,842** Amazon products restricted from shipping to at least one world region; **1.1%** of 796,081 sampled books restricted to at least one of four Middle Eastern countries {[knockel2026_banned]} | Common Crawl-derived Amazon product set | | | **17,842** Amazon products restricted from shipping to at least one world region; **1.1%** of 796,081 sampled books restricted to at least one of four Middle Eastern countries {[knockel2026_banned]} | Common Crawl-derived Amazon product set | |
| | Platform blocking of accounts advertised for sale worked on **19.71%** of them — YouTube 5.02%, Facebook 5.70%, X 18.67%, Instagram 46.41% {[beluri2025_exploration]} | 11,457 visible accounts from 11 marketplaces | | | Platform blocking of accounts advertised for sale worked on **19.71%** of them overall, and very unevenly: YouTube 5.02%, Facebook 5.70%, X 18.67%, Instagram 46.41%, TikTok 816 of 1,700 — the paper's own summary is "//TikTok and Instagram demonstrated the highest detection efficacy at 48%, whereas YouTube and Facebook showed the lowest efficacy at just 5%//" {[beluri2025_exploration]} | 11,457 visible accounts from 11 marketplaces | |
| | **136,009** Twitter users' Mastodon accounts identified across 2,879 instances; **96%** of migrants joined the largest quartile of instances {[he2023_flocking]} | 15,886 Mastodon instances, 1.02 M crawled accounts | | | **136,009** Twitter users' Mastodon accounts identified across 2,879 instances, and "//the top 25% most populous instances contain 96% of the users//" {[he2023_flocking]} | 15,886 Mastodon instances, 1.02 M crawled accounts | |
| | Over **90%** of targetable Facebook identities in the US had at least one data-broker-provided attribute (Australia 81.3%, UK 74.4%) {[venkatadri2019_auditing]} | the Facebook advertising interface, 7 countries | | | Over **90%** of targetable Facebook identities in the US had at least one data-broker-provided attribute (Australia 81.3%, UK 74.4%) {[venkatadri2019_auditing]} | the Facebook advertising interface, 7 countries | |
| | On X, posts containing external links had a median visibility score an **order of magnitude** below those without ("//0.0069 vs. 0.084//" for two named accounts) {[galeazzi2026_revealing]} | 17 M + 35 M tweets from two published datasets | | | On X, posts containing external links had a median visibility score an **order of magnitude** below those without ("//0.0069 vs. 0.084//" for two named accounts) {[galeazzi2026_revealing]} | 17 M + 35 M tweets from two published datasets | |
| | Sock-puppet accounts for personalised surfaces | **current, high-risk, unavoidable for feed studies** | 14 papers use the term; {[vombatkere2024_tiktok]}, {[karnam2026_setting]} | | | Sock-puppet accounts for personalised surfaces | **current, high-risk, unavoidable for feed studies** | 14 papers use the term; {[vombatkere2024_tiktok]}, {[karnam2026_setting]} | |
| | Logged-out scraping | **current, contested; a designated VLOP was fined for forbidding it** | 157 papers name a scraper; EC decision 5 December 2025 | | | Logged-out scraping | **current, contested; a designated VLOP was fined for forbidding it** | 157 papers name a scraper; EC decision 5 December 2025 | |
| | Third-party archives (Pushshift) | **current in practice, fragile** | 35 papers, still 9 in 2025 | | | Third-party archives (Pushshift) | **the live service is closed to researchers; the historical dumps are what the field is actually using** | 35 papers, still 9 in 2025 | |
| | Reuse of a published platform dataset | **current and dominant — often a symptom, not a choice** | 383 of 897 papers | | | Reuse of a published platform dataset | **current and very common — often a symptom, not a choice** | 383 of 897 name one; 208 have no primary collection at all | |
| | On-device / client-side model extraction | **current, niche, powerful** | {[west2024_picture]} | | | On-device / client-side model extraction | **current, niche, powerful** | {[west2024_picture]} | |
| | Amazon Product Advertising API 5.0 | **retired** | deprecated in favour of the Creators API; calls now return HTTP 403 ''AccessDeniedException''((https://affiliate-program.amazon.com/creatorsapi/docs/en-us/paapiv5-deprecation — fetched 2026-08-27.)) | | | Amazon Product Advertising API 5.0 | **retired** | deprecated in favour of the Creators API; calls now return HTTP 403 ''AccessDeniedException''((https://affiliate-program.amazon.com/creatorsapi/docs/en-us/paapiv5-deprecation — fetched 2026-08-27.)) | |
| |
| ^ Promised page ^ Corpus support ^ Decision ^ | ^ Promised page ^ Corpus support ^ Decision ^ |
| | **TikTok** | 7 candidate papers, **4** genuine after a full audit, none before 2023 | **No page.** Four papers is a paragraph, and it is above. | | | **TikTok** | 7 candidate papers, **4** genuine after a full audit, none before 2023 | **No page.** Not because TikTok is unimportant — because the TikTok measurement literature is mostly //outside these seven venues//, so a page built from this corpus would be four papers and would misrepresent the field. The four are named above. | |
| | **Amazon** | 72 candidates, **36%** precision; the genuine ones split between the Alexa voice/skill ecosystem ({[cheng2020_dangerous]}, {[liao2024_gdpr]}) and the retail storefront ({[knockel2026_banned]}) | **No page.** "Amazon" is not one measurement object, and the skill-store work belongs with [[Design:Mobile and app measurement]]. | | | **Amazon** | 72 candidates, **36%** precision; the genuine ones split between the Alexa voice/skill ecosystem ({[cheng2020_dangerous]}, {[liao2024_gdpr]}) and the retail storefront ({[knockel2026_banned]}) | **No page.** "Amazon" is not one measurement object, and the Alexa skill-store work is closer to app-store measurement than to this page, though [[Design:Mobile and app measurement]] does not cover skill stores yet. | |
| | **Twitter / X** | 142 candidates, **73%** precision — the largest coherent body | **No page.** The access route that produced nearly all of it no longer exists, so a page would be a history of a closed API. What survives is the access-route material above. | | | **Twitter / X** | 142 candidates, **73%** precision — the largest coherent body | **No page.** The access route that produced nearly all of it no longer exists, so a page would be a history of a closed API. What survives is the access-route material above. | |
| | **Facebook / Meta** | 151 candidates, **67%** precision | **No page.** The method-bearing parts are already elsewhere: the pixel and server-side flows on [[Privacy:Server side tracking]] and [[Privacy:Requests]], the SDKs on [[Design:Mobile and app measurement]], consent on [[Privacy:consent]]. The ad archive is the one genuinely distinct instrument, and it is not Facebook-specific. | | | **Facebook / Meta** | 151 candidates, **69%** precision | **No page.** The method-bearing parts are already elsewhere: the pixel and server-side flows on [[Privacy:Server side tracking]] and [[Privacy:Requests]], the SDKs on [[Design:Mobile and app measurement]], consent on [[Privacy:consent]]. The ad archive is the one genuinely distinct instrument, and it is not Facebook-specific. | |
| |
| A per-company page is the wrong axis anyway. What a platform study reuses across companies is the **route in** and the **denominator that route implies**; what does not transfer is the company's current API surface, which is precisely the part that goes stale fastest. This page is organised by route for that reason. | A per-company page is the wrong axis anyway. What a platform study reuses across companies is the **route in** and the **denominator that route implies**; what does not transfer is the company's current API surface, which is precisely the part that goes stale fastest. This page is organised by route for that reason. |
| * **What does a Meta Content Library study look like?** Zero examples in these venues. In particular: how do you report a denominator that is defined by a follower threshold, and can you publish anything from inside a controlled computing environment that satisfies an artefact badge? | * **What does a Meta Content Library study look like?** Zero examples in these venues. In particular: how do you report a denominator that is defined by a follower threshold, and can you publish anything from inside a controlled computing environment that satisfies an artefact badge? |
| * **What is the real cost of a metered-API study?** Nobody has published the arithmetic. A short note with a worked budget — reads, cap, and what you had to drop — would be worth more than another dataset. | * **What is the real cost of a metered-API study?** Nobody has published the arithmetic. A short note with a worked budget — reads, cap, and what you had to drop — would be worth more than another dataset. |
| * **Does the field's move to reused corpora change its findings?** 42.7% of platform papers already work from existing datasets. Whether the conclusions drawn from a 2019 Twitter corpus still describe X in 2026 is an answerable question nobody in this corpus asks. | * **Does the field's move to reused corpora change its findings?** 383 of 897 platform papers name an existing dataset and 208 collected nothing of their own. Whether the conclusions drawn from a 2019 Twitter corpus still describe X in 2026 is an answerable question nobody in this corpus asks. |
| | * **Search-engine and ads-ecosystem auditing has no page on this wiki**, and it is the largest measured family here — Google is rank 1 with 323 papers. SERP audits and ad-delivery audits share this page's problems (no API, sock puppets, personalisation) but have their own instruments. Out of scope here; worth its own page. |
| * **Ad-transparency archives** deserve their own page (see above). It needs the archive-by-archive coverage comparison that {[benzaamia2026_year]} starts. | * **Ad-transparency archives** deserve their own page (see above). It needs the archive-by-archive coverage comparison that {[benzaamia2026_year]} starts. |
| </WRAP> | </WRAP> |
| * **[[Design:User studies]]** — because "Amazon" on this page is 179 papers using **Mechanical Turk**. Crowdworkers labelling your platform data are annotation, not a user study. | * **[[Design:User studies]]** — because "Amazon" on this page is 179 papers using **Mechanical Turk**. Crowdworkers labelling your platform data are annotation, not a user study. |
| * **[[Design:Crawling location]]** — the vantage point, and why a datacenter IP gets a different platform than a residential one. | * **[[Design:Crawling location]]** — the vantage point, and why a datacenter IP gets a different platform than a residential one. |
| * **[[Design:Mobile and app measurement]]** — app stores, SDKs, and the on-device route {[west2024_picture]} used. | * **[[Design:Mobile and app measurement]]** — app stores, SDKs, static versus dynamic analysis, and certificate pinning. It does **not** currently cover voice-assistant skill stores or the on-device model-extraction route {[west2024_picture]} used; those are described here instead. |
| * **[[Design:Longitudinal]]** — repeating a platform measurement when the platform, not just the web, moved between waves. | * **[[Design:Longitudinal]]** — repeating a platform measurement when the platform, not just the web, moved between waves. |
| * **[[Programming:Registration]]** — the mechanics of the accounts a sock-puppet study needs. | * **[[Programming:Registration]]** — the mechanics of the accounts a sock-puppet study needs. |