User Tools

Site Tools


design:platforms

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
design:platforms [2026/08/27 20:31] – Currency review fixes: X pay-per-use launched 6 Feb 2026 and Basic/Pro remain for existing subscribers (the page wrongly said Basic was gone), Owned Reads 20 Apr 2026; Pushshift live access is moderator-only since 2023, not a researcher route; TikTok's EU karel.kubicek.claudedesign:platforms [2026/08/27 20:51] (current) – Quote the papers' own contiguous sentences for the Mastodon 96% and the WhatsApp candidate-number count (both were paraphrases of column-spliced text). Authored by Claude karel.kubicek.claude
Line 12: Line 12:
 Meanwhile the official researcher routes that replaced the open APIs are **almost entirely absent from this literature**. Across the **5,855** papers with full text: the **Meta Content Library** appears **once**, and that once is a parenthetical noting that CrowdTangle "//now replaced by the Meta Content Library//" {[roy2025_darkgram]}. The **TikTok Research API** appears in **zero** papers. The **YouTube Researcher Program** appears in **two**. The X/Twitter **Academic Research track** appears in **eight**. By contrast **Pushshift** — a third-party Reddit archive — appears in **35**, and **Ad Library** in **69**. Meanwhile the official researcher routes that replaced the open APIs are **almost entirely absent from this literature**. Across the **5,855** papers with full text: the **Meta Content Library** appears **once**, and that once is a parenthetical noting that CrowdTangle "//now replaced by the Meta Content Library//" {[roy2025_darkgram]}. The **TikTok Research API** appears in **zero** papers. The **YouTube Researcher Program** appears in **two**. The X/Twitter **Academic Research track** appears in **eight**. By contrast **Pushshift** — a third-party Reddit archive — appears in **35**, and **Ad Library** in **69**.
  
-If you plan a platform study in 2026 through an official research programme, assume you are an early adopter: there is no methods section in these venues to copy, and the application lead time is weeks, not days.+Part of that absence is **mechanical**, and the page says so rather than letting you infer otherwise: these programmes are 2023-and-later launches read through a corpus whose 2025–2026 slice is under-covered by construction, and the venues that publish most platform work of this kind — ICWSM, CHI, FAccT, the communications journals — are **not in the corpus at all**. The right reading is "there is no worked example //in these seven venues//", not "the programmes do not work"
 + 
 +Either way the consequence for you is the same. If you plan a platform study in 2026 through an official research programme, assume you are an early adopter: there is no methods section in these venues to copy, and the application lead time is weeks, not days.
 </WRAP> </WRAP>
  
Line 22: Line 24:
   * {[vombatkere2024_tiktok]} (TheWebConf 2024) — **sock-puppet accounts** on a feed you cannot query. Five bot accounts plus a donated dataset of 347 real users; concludes that TikTok "//exploits real users' interests in between 30% and 50% of all recommended videos in the first thousand videos//". Read it for how you measure a recommender with no API at all.   * {[vombatkere2024_tiktok]} (TheWebConf 2024) — **sock-puppet accounts** on a feed you cannot query. Five bot accounts plus a donated dataset of 347 real users; concludes that TikTok "//exploits real users' interests in between 30% and 50% of all recommended videos in the first thousand videos//". Read it for how you measure a recommender with no API at all.
   * {[karnam2026_setting]} (IEEE S&P 2026) — the **data subject's right of access** as the instrument. Sock-puppet accounts on Instagram, TikTok and YouTube, then GDPR Art. 15 requests, then a comparison of the returned data package against the ground truth the authors themselves generated. All three platforms failed to report purpose, recipients and retention period.   * {[karnam2026_setting]} (IEEE S&P 2026) — the **data subject's right of access** as the instrument. Sock-puppet accounts on Instagram, TikTok and YouTube, then GDPR Art. 15 requests, then a comparison of the returned data package against the ground truth the authors themselves generated. All three platforms failed to report purpose, recipients and retention period.
-  * {[gegenhuber2026_there]} (NDSS 2026) — **enumeration at platform scale**, and the ethics that go with it. Probes 63 billion candidate phone numbers and discovers **3,546,479,731** WhatsApp accounts across 245 countries, **57%** of them with a public profile picture. Read it for what a platform-scale study has to say about disclosure and harm before a programme committee will accept it.+  * {[gegenhuber2026_there]} (NDSS 2026) — **enumeration at platform scale**, and the ethics that go with it. Probes 63,170,000,000 candidate phone numbers and discovers **3,546,479,731** WhatsApp accounts across 245 countries, **57%** of them with a public profile picture. Read it for what a platform-scale study has to say about disclosure and harm before a programme committee will accept it.
  
 If you only read one thing about the **rules**, read the Commission's December 2025 decision fining X €120 million: it is the first DSA non-compliance decision, and one of the three grounds is research data access.((European Commission, //Commission fines X €120 million under the Digital Services Act//, press release, 5 December 2025, https://digital-strategy.ec.europa.eu/en/news/commission-fines-x-eu120-million-under-digital-services-act — page fetched 2026-08-27, "Last update 16 January 2026".)) If you only read one thing about the **rules**, read the Commission's December 2025 decision fining X €120 million: it is the first DSA non-compliance decision, and one of the three grounds is research data access.((European Commission, //Commission fines X €120 million under the Digital Services Act//, press release, 5 December 2025, https://digital-strategy.ec.europa.eu/en/news/commission-fines-x-eu120-million-under-digital-services-act — page fetched 2026-08-27, "Last update 16 January 2026".))
Line 47: Line 49:
 Every route below is real and in use somewhere. What differs is whether it is current, what it costs, and what population it yields. Dates are as of **2026-08-27** and verified against primary sources; the [[provenance:design:platforms|provenance page]] lists each URL, quote and fetch date, and names the claims that could not be verified. Every route below is real and in use somewhere. What differs is whether it is current, what it costs, and what population it yields. Dates are as of **2026-08-27** and verified against primary sources; the [[provenance:design:platforms|provenance page]] lists each URL, quote and fetch date, and names the claims that could not be verified.
  
-^ Route ^ Status in 2026 ^ What it gets you ^ Papers in our corpus that name it ^ +The **Status** column is one word; the dated evidence for each verdict is in [[#Which Methods Are Current]] at the foot of the page. 
-| **Open public API, free** | **historical** for the platforms this page is about | the population the platform chose to expose | 91 papers name a Twitter/X streaming or search API, peaking 2022 | + 
-| **Paid metered API** | **current** — X moved to pay-per-usage | whatever you can afford | see the arithmetic below | +^ Route ^ Status ^ What it gets you ^ Papers in our corpus that name it ^ 
-| **Platform research programme** | **current, and unused in this literature** | a curated, filtered archive inside a controlled environment | Meta Content Library **1**, TikTok Research API **0**, YouTube Researcher Program **2** | +| **Open public API, free** | **historical** | the population the platform chose to expose | 91 papers name a Twitter/X streaming or search API, peaking 2022 | 
-| **Regulator route (DSA Art. 40 vetted researcher)** | **current, newly operational** | in principle non-public data from a designated VLOP | effectively **zero** papers; see the audit below |+| **Paid metered API** | **current** | whatever you can afford | X moved to pay-per-usagesee the arithmetic below | 
 +| **Platform research programme** | **current** | a curated, filtered archive inside a controlled environment | Meta Content Library **1**, TikTok Research API **0**, YouTube Researcher Program **2** | 
 +| **Regulator route (DSA Art. 40 vetted researcher)** | **current** | in principle non-public data from a designated VLOP | effectively **zero** papers; see the audit below |
 | **Transparency / ad archive** | **current** | only what the archive covers — usually ads that actually ran | **69** papers name an Ad Library or ad archive | | **Transparency / ad archive** | **current** | only what the archive covers — usually ads that actually ran | **69** papers name an Ad Library or ad archive |
-| **Third-party archive of platform data** | **historical for researchers, but its dumps still circulate** | someone else's crawl, with their gaps | **35** papers name Pushshift | +| **Third-party archive of platform data** | **historical** | someone else's crawl, with their gaps — the dumps circulate, the live service does not | **35** papers name Pushshift | 
-| **Reuse of a published platform dataset** | **current, and very common** | a frozen snapshot you did not design | **383 of 897** name an existing dataset; **208** name one //and no primary collection//+| **Reuse of a published platform dataset** | **current** | a frozen snapshot you did not design | **383 of 897** name an existing dataset; **208** name one //and no primary collection//
-| **Logged-out scraping** | **current, legally contested — and now the subject of a DSA fine against a platform that forbade it** | the public surface, as a bot sees it | **157** papers name a scraping or automation tool | +| **Logged-out scraping** | **current, contested** | the public surface, as a bot sees it | **157** papers name a scraping or automation tool | 
-| **Accounts you created (sock puppets)** | **current, highest-risk** | the logged-in surface, and anything personalised | **14** papers use the term "sock puppet"; **22** state `account-registration` | +| **Accounts you created (sock puppets)** | **current** | the logged-in surface, and anything personalised | **14** papers use the term "sock puppet"; **22** are labelled `account-registration` | 
-| **Data donation** | **current and rising** | real users' own data, with consent | **25** papers mention data donation, **11** of them in 2025 | +| **Data donation** | **current** | real users' own data, with consent | **25** papers mention data donation, **11** of them in 2025 | 
-| **Right of access / DSAR** | **current and rising** | what the platform says it holds about a subject | **52** papers, rising from 2022 | +| **Right of access / DSAR** | **current** | what the platform says it holds about a subject | **52** papers, rising from 2022 | 
-| **On-device or client-side extraction** | **niche, powerful** | the model or logic actually shipped to users | {[west2024_picture]} extracted the on-device ML models from the Instagram and TikTok apps | +| **On-device or client-side extraction** | **niche** | the model or logic actually shipped to users | {[west2024_picture]} extracted the on-device ML models from the Instagram and TikTok apps | 
-| **Operator collaboration** | **current, not replicable** | everything, and no reproducibility | {[cohn2020_delf]}, {[schlinker2019_internet]} |+| **Operator collaboration** | **not replicable** | everything, and no reproducibility | {[cohn2020_delf]}, {[schlinker2019_internet]} |
  
 ==== The metered API changes the study design, not just the budget ==== ==== The metered API changes the study design, not just the budget ====
  
-X's API is now **pay-per-usage with no subscription**: %%$0.005%% per Post read, %%$0.010%% per User read, and "//Pay-per-usage plans are capped at 3 million Post reads per monthly billing cycle//" with Enterprise above that.((https://docs.x.com/x-api/getting-started/pricing — fetched 2026-08-27; the page carries no date. There is no free tier and no academic tier in the current documentation tree.)) One million posts is therefore **%%$5,000%%** and the monthly ceiling on the self-serve route is **3 million posts**. A study that used to be "collect the 1% stream for six months" is now a budget line and a cap.+X's API is now **pay-per-usage with no subscription**: %%$0.005%% per Post read, %%$0.010%% per User read, and "//Pay-per-usage plans are capped at 3 million Post reads per monthly billing cycle//" with Enterprise above that.((https://docs.x.com/x-api/getting-started/pricing — fetched 2026-08-27; the page carries no date. On the same fetch, the //About the X API// page lists only v2 (pay-per-usage) and v1.1 ("limited support"), and neither it nor the pricing page mentions a free or academic tier; we did not exhaustively search the whole documentation tree, so read that as "not offered on the pages a new applicant is sent to" rather than as a proof of absence.)) One million posts is therefore **%%$5,000%%** and the monthly ceiling on the self-serve route is **3 million posts**. A study that used to be "collect the 1% stream for six months" is now a budget line and a cap.
  
 This is not hypothetical, and the field has already adapted. {[nguyen2025_please]} (USENIX Security 2025) built its collection pipeline around the read cap: it polls a **counts** endpoint, which "//does not add to our monthly tweet read limit//", and only fetches actual posts when the count is non-zero — because "//the total number of tweets that can be downloaded is very small for the Basic tier//". That is still a sound pattern, but the tier it names is no longer what a new project gets: X's own changelog records that pay-per-usage "//officially launched//" on **6 February 2026** as the self-serve model, with the older Basic and Pro plans "//remain[ing] available//" only to existing subscribers, and then a further price change on **20 April 2026** introducing "//Owned Reads//" at %%$0.001%% per resource for a developer's own account data.((https://docs.x.com/changelog — fetched 2026-08-27; entries dated 6 February 2026 ("Launch of X API Pay-Per-Use pricing") and 16 April 2026 ("X API pricing update: Owned Reads now %%$0.001%%").)) The pricing model of the single most-measured platform in this corpus changed **twice inside 2026**. Date it, and re-check it before you submit. This is not hypothetical, and the field has already adapted. {[nguyen2025_please]} (USENIX Security 2025) built its collection pipeline around the read cap: it polls a **counts** endpoint, which "//does not add to our monthly tweet read limit//", and only fetches actual posts when the count is non-zero — because "//the total number of tweets that can be downloaded is very small for the Basic tier//". That is still a sound pattern, but the tier it names is no longer what a new project gets: X's own changelog records that pay-per-usage "//officially launched//" on **6 February 2026** as the self-serve model, with the older Basic and Pro plans "//remain[ing] available//" only to existing subscribers, and then a further price change on **20 April 2026** introducing "//Owned Reads//" at %%$0.001%% per resource for a developer's own account data.((https://docs.x.com/changelog — fetched 2026-08-27; entries dated 6 February 2026 ("Launch of X API Pay-Per-Use pricing") and 16 April 2026 ("X API pricing update: Owned Reads now %%$0.001%%").)) The pricing model of the single most-measured platform in this corpus changed **twice inside 2026**. Date it, and re-check it before you submit.
Line 104: Line 108:
   * **A research archive is not the platform.** The Meta Content Library includes Facebook posts to Pages, groups and events "//as well as posts that appear on public profiles that are either verified or that have 100 or more followers//"; Instagram business and creator accounts plus personal accounts "//verified or that have 100 or more followers//"; Threads public profiles with 100 or more followers.((https://transparency.meta.com/researchtools/meta-content-library/, fetched 2026-08-27.)) A prevalence computed there is a prevalence **among accounts above a follower threshold**, not among users.   * **A research archive is not the platform.** The Meta Content Library includes Facebook posts to Pages, groups and events "//as well as posts that appear on public profiles that are either verified or that have 100 or more followers//"; Instagram business and creator accounts plus personal accounts "//verified or that have 100 or more followers//"; Threads public profiles with 100 or more followers.((https://transparency.meta.com/researchtools/meta-content-library/, fetched 2026-08-27.)) A prevalence computed there is a prevalence **among accounts above a follower threshold**, not among users.
   * **An ad archive contains ads that ran.** It cannot tell you what was rejected, and it can be wrong in both directions about what is political — which is exactly {[bouchaud2024_beyond]}'s finding, on both sides at once.   * **An ad archive contains ads that ran.** It cannot tell you what was rejected, and it can be wrong in both directions about what is political — which is exactly {[bouchaud2024_beyond]}'s finding, on both sides at once.
-  * **A sampled stream is a sample.** {[hagen2021_numbers]} is explicit that its dataset came from "//the 1% streaming API that Twitter provides to vetted researchers//", so all of its figures are lower bounds. Say which stream and which sampling rate.+  * **A sampled stream is a sample.** {[kaleli2021_human]} is explicit that its dataset came from "//the 1% streaming API that Twitter provides to vetted researchers//" and that consequently "//all the numbers that we presented in this paper are lower bounds//". Say which stream and which sampling rate, and say which direction the bias runs. 
 +  * **We could not establish what Reddit currently offers a researcher.** Every Reddit-owned domain refused our fetches, so this page makes **no claim** about Reddit's API pricing, rate limits or researcher terms — only about what the corpus shows (63 papers measuring Reddit, 14 naming a Reddit API or PRAW, 35 naming Pushshift). Check it yourself before planning around it, and see [[provenance:design:platforms]] for exactly what failed.
   * **A third-party archive has the archiver's gaps, not yours.** Pushshift is still the most-named Reddit source in this corpus (**35** papers, e.g. {[zannettou2018_origins]}), and its coverage window and completeness are properties of Pushshift, not of Reddit. It is also **no longer a route you can open**: Reddit's own moderator help page says Pushshift API access "//will be reinstated for verified Reddit moderators//", that each one "//need[s] explicit approval from Reddit//", and that "//the use of Pushshift will be limited to moderation use cases only//".((https://support.reddithelp.com/hc/en-us/articles/16470271632404-Pushshift-Access-Request — fetched 2026-08-27 with a headless browser (the page is Cloudflare-walled to ''%%curl%%''); the page is stamped "Updated 1 year ago". A paper citing Pushshift today is citing a historical dump, and should say which one and when it was obtained.))   * **A third-party archive has the archiver's gaps, not yours.** Pushshift is still the most-named Reddit source in this corpus (**35** papers, e.g. {[zannettou2018_origins]}), and its coverage window and completeness are properties of Pushshift, not of Reddit. It is also **no longer a route you can open**: Reddit's own moderator help page says Pushshift API access "//will be reinstated for verified Reddit moderators//", that each one "//need[s] explicit approval from Reddit//", and that "//the use of Pushshift will be limited to moderation use cases only//".((https://support.reddithelp.com/hc/en-us/articles/16470271632404-Pushshift-Access-Request — fetched 2026-08-27 with a headless browser (the page is Cloudflare-walled to ''%%curl%%''); the page is stamped "Updated 1 year ago". A paper citing Pushshift today is citing a historical dump, and should say which one and when it was obtained.))
   * **A reused dataset freezes a platform that has since changed.** **383 of the 897** platform-subject papers name an existing dataset (''temporal.mode'' is multi-valued, so 175 of those also collected something themselves; **208** collected nothing). If that is you, date the snapshot and say what changed on the platform since — the four false positives in our Twitter/X audit were all papers using Twitter-derived benchmark corpora with no relationship to the platform at collection time.   * **A reused dataset freezes a platform that has since changed.** **383 of the 897** platform-subject papers name an existing dataset (''temporal.mode'' is multi-valued, so 175 of those also collected something themselves; **208** collected nothing). If that is you, date the snapshot and say what changed on the platform since — the four false positives in our Twitter/X audit were all papers using Twitter-derived benchmark corpora with no relationship to the platform at collection time.
-  * **A cross-platform study has one denominator per platform, and they are not comparable.** {[acharya2024_imitation]} looked for brand-impersonation accounts targeting the top 10K Tranco brands on X, Instagram, Telegram and YouTube at once; {[beluri2025_exploration]} tracked accounts advertised for sale across five platforms and found blocking efficacy ranging from **5.02%** (YouTube) to **46.41%** (Instagram). A cross-platform rate is only meaningful if you say what you could see on each platform, because the routes in differ and so does the visible surface. A gap between two platforms may be a gap between two access routes. +  * **A cross-platform study has one denominator per platform, and they are not comparable.** {[acharya2024_imitation]} looked for brand-impersonation accounts targeting the top 10K Tranco brands on X, Instagram, Telegram and YouTube at once; {[beluri2025_exploration]} tracked accounts advertised for sale across five platforms and found blocking efficacy ranging from **5.02%** (YouTube) up to **48%** — its own summary is that "//TikTok and Instagram demonstrated the highest detection efficacy at 48%, whereas YouTube and Facebook showed the lowest efficacy at just 5%//". A cross-platform rate is only meaningful if you say what you could see on each platform, because the routes in differ and so does the visible surface. A gap between two platforms may be a gap between two access routes. 
-  * **The logged-out surface is not the platform either.** Of the **260** platform-subject papers that report a crawl configuration, **145 (55.8%)** state that they used **no authentication at all**, only **22** registered an account, **4** logged in manually, and **none** used automated login or SSO. Whatever those studies measured, it is what an anonymous visitor sees.+  * **The logged-out surface is not the platform either.** Of the **260** platform-subject papers that report a crawl configuration, **145 (55.8%)** are labelled as using **no authentication at all**, **22** as registering an account, **4** as logging in manually, and **none** as using automated login or SSO. Whatever those studies measured, it is mostly what an anonymous visitor sees. **These are schema labels, not audited ones**, and [[Programming:Registration]] is explicit about why that matters for this exact field: ''%%crawlConfig%%'' carries **one evidence quote for the whole object**, so the ''%%authentication%%'' label cannot be checked against its quote, and a hand-audit of **another field of that same object**, ''%%consentAction%%'', found 19.4% false positives ([[Privacy:consent]]). We did not audit these 145. Treat the shape as real and the precise share as unverified.
  
 ===== Rate Limits, Quotas, Bans and the Login Wall ===== ===== Rate Limits, Quotas, Bans and the Login Wall =====
Line 123: Line 128:
 | sock puppet | 14 (**1.6%**) | 26 (0.4%) | 3 of 108 (2.8%) | | sock puppet | 14 (**1.6%**) | 26 (0.4%) | 3 of 108 (2.8%) |
  
-Read the ToS row as the headline: a platform paper is roughly **twice as likely** as an average paper in these venues to discuss terms of service, and in 2025 more than one in five did. That is the direction the reviewing is moving.+Read the ToS row as the headline: at **14.9%**, a platform paper is roughly **twice as likely as the average paper in these seven venues** (7.8%) to discuss terms of service, and in 2025 more than one in five did. Against a narrower and fairer baseline — papers that ran a crawl, where [[Programming:Registration]] reports 11.5% — the ratio is about 1.3×. Either way the direction is the same, and it is the direction the reviewing is moving.
  
-Three practical consequences:+Four practical consequences:
  
   * **Design for the limit, not around it.** {[nguyen2025_please]}'s counts-endpoint trigger is the pattern: find the cheap query that tells you whether the expensive query is worth making. Report the limit you worked under, because it bounds your recall.   * **Design for the limit, not around it.** {[nguyen2025_please]}'s counts-endpoint trigger is the pattern: find the cheap query that tells you whether the expensive query is worth making. Report the limit you worked under, because it bounds your recall.
   * **The limit can be the finding.** {[xue2021_throttling]} measured Russian ISPs throttling Twitter and showed the mechanism was SNI-based: "//throttling is triggered upon observing Twitter-related domains (*.twimg.com, twitter.com, t.co) in the SNI//". If you are being throttled and can characterise by what, that is a result.   * **The limit can be the finding.** {[xue2021_throttling]} measured Russian ISPs throttling Twitter and showed the mechanism was SNI-based: "//throttling is triggered upon observing Twitter-related domains (*.twimg.com, twitter.com, t.co) in the SNI//". If you are being throttled and can characterise by what, that is a result.
   * **Expect the platform to fight your crawler, and say whether it did.** Only **14 of 260** platform-subject papers with a crawl configuration (**5.4%**) say anything about robots.txt — barely above the **4.7%** baseline for all papers with a crawl configuration. If the platform served you a bot challenge, that is a measurement result about the platform, not an embarrassment.   * **Expect the platform to fight your crawler, and say whether it did.** Only **14 of 260** platform-subject papers with a crawl configuration (**5.4%**) say anything about robots.txt — barely above the **4.7%** baseline for all papers with a crawl configuration. If the platform served you a bot challenge, that is a measurement result about the platform, not an embarrassment.
-  * **Sock puppets are the highest-risk route.** They get you the personalised surface — {[vombatkere2024_tiktok]} and {[karnam2026_setting]} both need them — and they are the fact pattern most likely to breach terms of service and, in the US, to have been litigated. Get ethics review first (see [[Practices:Ethics]]), keep the accounts minimal, and write down what you did: **only 32.5%** of the **841** empirical platform-subject papers state an ethics-review outcome, statistically indistinguishable from the **33.8%** corpus baseline. That is not good enough for a study whose method is creating fake accounts on someone else's service.+  * **Sock puppets are the highest-risk route.** They get you the personalised surface — {[vombatkere2024_tiktok]} and {[karnam2026_setting]} both need them — and they are the fact pattern most likely to breach terms of service and, in the US, to have been litigated. Get ethics review first (see [[Practices:Ethics]]), keep the accounts minimal, and write down what you did: **only 32.5%** of the **841** empirical platform-subject papers state an ethics-review outcome — essentially at the **33.8%** corpus baseline (no test was run, and the platform papers are inside that baseline). That is not good enough for a study whose method is creating fake accounts on someone else's service.
  
 ==== Terms of service are no longer only the platform's weapon ==== ==== Terms of service are no longer only the platform's weapon ====
Line 154: Line 159:
 | Amazon | 72 | 11 (every 7th) | 4 | **36%** | Amazon Reviews benchmarks, devices //bought on// Amazon, a hosted speech service | | Amazon | 72 | 11 (every 7th) | 4 | **36%** | Amazon Reviews benchmarks, devices //bought on// Amazon, a hosted speech service |
  
-Treat the ranking below as a ranking. Do not quote a family's count as a precise number of papers without the precision above attached to it.+Treat the ranking below as a ranking. Do not quote a family's count as a precise number of papers without the precision above attached to it — and note that these samples are **small**: at //n// = 15 the 95% interval around 73% is roughly ±20 points. They are coarse corrections, not measurements.
  
 ==== Which platforms the field measures ==== ==== Which platforms the field measures ====
Line 181: Line 186:
 | 19 | eBay | 6 | 0.1% | | 19 | eBay | 6 | 0.1% |
  
-The top two rows are largely **app-store** work, which belongs to [[Design:Mobile and app measurement]] rather than here. Note the last rows: **TikTok is 7 papers**, of which 4 survived a full audit. There is no body of TikTok measurement in these seven venues to systematise.+Rows 1 and 7 — Google and the Apple App Store — are largely **app-store** work, which is closer to [[Design:Mobile and app measurement]] than to this page. Note also what this page does **not** cover: search-engine result auditing and the Google ads ecosystem are a real body of work inside these venues and have no page on this wiki yet; they are out of scope here and named in [[#Open Questions]]. Note the last rows: **TikTok is 7 papers**, of which 4 survived a full audit. There is no body of TikTok measurement in these seven venues to systematise.
  
 ==== Amazon means five different things, and mostly not the shop ==== ==== Amazon means five different things, and mostly not the shop ====
Line 195: Line 200:
  
 A substring match on "Amazon" therefore over-counts Amazon platform measurement by roughly **sixteen to one**. An earlier iteration of our own script did exactly this and reported 176 Amazon papers. A substring match on "Amazon" therefore over-counts Amazon platform measurement by roughly **sixteen to one**. An earlier iteration of our own script did exactly this and reported 176 Amazon papers.
 +
 +<WRAP tip>
 +**Two other pages count these same names and get different numbers. Both are right; they answer different questions.** [[Design:Website selection]] reports **463** papers for Alexa, folded over the 1,143 that drew a **web-unit** population. The **412** above is the subset our role rule assigns to the ranking list, over all 5,492 papers with a stated population source. Reconciling them: **486** papers name "Alexa" in a stated population source at all, **464** of those also have a web-unit population tuple — which is the figure [[Design:Website selection]] is reporting — and 412 is what survives after Alexa-as-skill-store strings are diverted to the ''%%subject%%'' role. Likewise our **179** Mechanical Turk papers is a count of ''%%population.sourceList%%'' strings; [[Design:User studies]] publishes three //different// bounds on the same thing — 139 by the schema enum, 92 by quote, 279 by full text — and warns against reading any of them as usage. Neither fold has been reconciled with the other, and this page does not claim its number is the better one.
 +</WRAP>
  
 ==== The Twitter/X curve, and what it does and does not show ==== ==== The Twitter/X curve, and what it does and does not show ====
Line 231: Line 240:
 | Meta / Facebook Ad Library | 5 | 0.6% | | Meta / Facebook Ad Library | 5 | 0.6% |
 | CrowdTangle | 5 | 0.6% | | CrowdTangle | 5 | 0.6% |
-TikTok API / TikTok-Api | 1 | 0.1% |+''%%TikTok-Api%%'' — the **unofficial** scraper library, not the Research API | 1 | 0.1% |
 | **named no tool for it at all** | **552** | **61.5%** | | **named no tool for it at all** | **552** | **61.5%** |
  
-That last row is the reporting gap on this page: **61.5%** of platform-subject papers name no instrument by which they reached the platformThe route is the design decision, and three papers in five do not state it.+That last row is the reporting gap on this page: for **61.5%** of platform-subject papers, **no named instrument survives into this corpus's tool lists**Read that as an upper bound on the gap rather than as "three in five do not state it": extraction recall on free-text tool names is imperfect, and we did not hand-audit the 552. Even as an upper bound it is the largest single reporting hole the page found, because the route is the design decision.
  
 ==== The reproducibility cost ==== ==== The reproducibility cost ====
Line 273: Line 282:
 ^ Finding ^ Denominator the paper used ^ ^ Finding ^ Denominator the paper used ^
 | Only **7.7%** of undeclared political ads were moderated as political by Meta; **60.4%** of the ads Meta did moderate did not match its own criteria {[bouchaud2024_beyond]} | 29.5 M ads from the Meta Ad Library API, 16 EU countries | | Only **7.7%** of undeclared political ads were moderated as political by Meta; **60.4%** of the ads Meta did moderate did not match its own criteria {[bouchaud2024_beyond]} | 29.5 M ads from the Meta Ad Library API, 16 EU countries |
-| **3,546,479,731** WhatsApp accounts discovered; **57%** with a public profile picture, **66%** of a 500,000-image sample containing a detectable face {[gegenhuber2026_there]} | 63.2 bn candidate phone numbers enumerated |+| **3,546,479,731** WhatsApp accounts discovered; **57%** with a public profile picture, **66%** of a 500,000-image sample containing a detectable face {[gegenhuber2026_there]} | 63,170,000,000 candidate phone numbers enumerated |
 | TikTok "//exploits real users' interests in between 30% and 50% of all recommended videos in the first thousand videos//" {[vombatkere2024_tiktok]} | 347 donating users, 4.9 M videos, plus 5 bot accounts | | TikTok "//exploits real users' interests in between 30% and 50% of all recommended videos in the first thousand videos//" {[vombatkere2024_tiktok]} | 347 donating users, 4.9 M videos, plus 5 bot accounts |
 | **17,842** Amazon products restricted from shipping to at least one world region; **1.1%** of 796,081 sampled books restricted to at least one of four Middle Eastern countries {[knockel2026_banned]} | Common Crawl-derived Amazon product set | | **17,842** Amazon products restricted from shipping to at least one world region; **1.1%** of 796,081 sampled books restricted to at least one of four Middle Eastern countries {[knockel2026_banned]} | Common Crawl-derived Amazon product set |
-| Platform blocking of accounts advertised for sale worked on **19.71%** of them — YouTube 5.02%, Facebook 5.70%, X 18.67%, Instagram 46.41% {[beluri2025_exploration]} | 11,457 visible accounts from 11 marketplaces | +| Platform blocking of accounts advertised for sale worked on **19.71%** of them overall, and very unevenly: YouTube 5.02%, Facebook 5.70%, X 18.67%, Instagram 46.41%, TikTok 816 of 1,700 — the paper's own summary is "//TikTok and Instagram demonstrated the highest detection efficacy at 48%, whereas YouTube and Facebook showed the lowest efficacy at just 5%//" {[beluri2025_exploration]} | 11,457 visible accounts from 11 marketplaces | 
-| **136,009** Twitter users' Mastodon accounts identified across 2,879 instances; **96%** of migrants joined the largest quartile of instances {[he2023_flocking]} | 15,886 Mastodon instances, 1.02 M crawled accounts |+| **136,009** Twitter users' Mastodon accounts identified across 2,879 instances, and "//the top 25% most populous instances contain 96% of the users//" {[he2023_flocking]} | 15,886 Mastodon instances, 1.02 M crawled accounts |
 | Over **90%** of targetable Facebook identities in the US had at least one data-broker-provided attribute (Australia 81.3%, UK 74.4%) {[venkatadri2019_auditing]} | the Facebook advertising interface, 7 countries | | Over **90%** of targetable Facebook identities in the US had at least one data-broker-provided attribute (Australia 81.3%, UK 74.4%) {[venkatadri2019_auditing]} | the Facebook advertising interface, 7 countries |
 | On X, posts containing external links had a median visibility score an **order of magnitude** below those without ("//0.0069 vs. 0.084//" for two named accounts) {[galeazzi2026_revealing]} | 17 M + 35 M tweets from two published datasets | | On X, posts containing external links had a median visibility score an **order of magnitude** below those without ("//0.0069 vs. 0.084//" for two named accounts) {[galeazzi2026_revealing]} | 17 M + 35 M tweets from two published datasets |
Line 323: Line 332:
  
 ^ Promised page ^ Corpus support ^ Decision ^ ^ Promised page ^ Corpus support ^ Decision ^
-| **TikTok** | 7 candidate papers, **4** genuine after a full audit, none before 2023 | **No page.** Four papers is a paragraph, and it is above. | +| **TikTok** | 7 candidate papers, **4** genuine after a full audit, none before 2023 | **No page.** Not because TikTok is unimportant — because the TikTok measurement literature is mostly //outside these seven venues//so a page built from this corpus would be four papers and would misrepresent the field. The four are named above. | 
-| **Amazon** | 72 candidates, **36%** precision; the genuine ones split between the Alexa voice/skill ecosystem ({[cheng2020_dangerous]}, {[liao2024_gdpr]}) and the retail storefront ({[knockel2026_banned]}) | **No page.** "Amazon" is not one measurement object, and the skill-store work belongs with [[Design:Mobile and app measurement]]. |+| **Amazon** | 72 candidates, **36%** precision; the genuine ones split between the Alexa voice/skill ecosystem ({[cheng2020_dangerous]}, {[liao2024_gdpr]}) and the retail storefront ({[knockel2026_banned]}) | **No page.** "Amazon" is not one measurement object, and the Alexa skill-store work is closer to app-store measurement than to this page, though [[Design:Mobile and app measurement]] does not cover skill stores yet. |
 | **Twitter / X** | 142 candidates, **73%** precision — the largest coherent body | **No page.** The access route that produced nearly all of it no longer exists, so a page would be a history of a closed API. What survives is the access-route material above. | | **Twitter / X** | 142 candidates, **73%** precision — the largest coherent body | **No page.** The access route that produced nearly all of it no longer exists, so a page would be a history of a closed API. What survives is the access-route material above. |
 | **Facebook / Meta** | 151 candidates, **69%** precision | **No page.** The method-bearing parts are already elsewhere: the pixel and server-side flows on [[Privacy:Server side tracking]] and [[Privacy:Requests]], the SDKs on [[Design:Mobile and app measurement]], consent on [[Privacy:consent]]. The ad archive is the one genuinely distinct instrument, and it is not Facebook-specific. | | **Facebook / Meta** | 151 candidates, **69%** precision | **No page.** The method-bearing parts are already elsewhere: the pixel and server-side flows on [[Privacy:Server side tracking]] and [[Privacy:Requests]], the SDKs on [[Design:Mobile and app measurement]], consent on [[Privacy:consent]]. The ad archive is the one genuinely distinct instrument, and it is not Facebook-specific. |
Line 339: Line 348:
   * **What is the real cost of a metered-API study?** Nobody has published the arithmetic. A short note with a worked budget — reads, cap, and what you had to drop — would be worth more than another dataset.   * **What is the real cost of a metered-API study?** Nobody has published the arithmetic. A short note with a worked budget — reads, cap, and what you had to drop — would be worth more than another dataset.
   * **Does the field's move to reused corpora change its findings?** 383 of 897 platform papers name an existing dataset and 208 collected nothing of their own. Whether the conclusions drawn from a 2019 Twitter corpus still describe X in 2026 is an answerable question nobody in this corpus asks.   * **Does the field's move to reused corpora change its findings?** 383 of 897 platform papers name an existing dataset and 208 collected nothing of their own. Whether the conclusions drawn from a 2019 Twitter corpus still describe X in 2026 is an answerable question nobody in this corpus asks.
 +  * **Search-engine and ads-ecosystem auditing has no page on this wiki**, and it is the largest measured family here — Google is rank 1 with 323 papers. SERP audits and ad-delivery audits share this page's problems (no API, sock puppets, personalisation) but have their own instruments. Out of scope here; worth its own page.
   * **Ad-transparency archives** deserve their own page (see above). It needs the archive-by-archive coverage comparison that {[benzaamia2026_year]} starts.   * **Ad-transparency archives** deserve their own page (see above). It needs the archive-by-archive coverage comparison that {[benzaamia2026_year]} starts.
 </WRAP> </WRAP>
Line 348: Line 358:
   * **[[Design:User studies]]** — because "Amazon" on this page is 179 papers using **Mechanical Turk**. Crowdworkers labelling your platform data are annotation, not a user study.   * **[[Design:User studies]]** — because "Amazon" on this page is 179 papers using **Mechanical Turk**. Crowdworkers labelling your platform data are annotation, not a user study.
   * **[[Design:Crawling location]]** — the vantage point, and why a datacenter IP gets a different platform than a residential one.   * **[[Design:Crawling location]]** — the vantage point, and why a datacenter IP gets a different platform than a residential one.
-  * **[[Design:Mobile and app measurement]]** — app stores, SDKs, and the on-device route {[west2024_picture]} used.+  * **[[Design:Mobile and app measurement]]** — app stores, SDKs, static versus dynamic analysis, and certificate pinning. It does **not** currently cover voice-assistant skill stores or the on-device model-extraction route {[west2024_picture]} used; those are described here instead.
   * **[[Design:Longitudinal]]** — repeating a platform measurement when the platform, not just the web, moved between waves.   * **[[Design:Longitudinal]]** — repeating a platform measurement when the platform, not just the web, moved between waves.
   * **[[Programming:Registration]]** — the mechanics of the accounts a sock-puppet study needs.   * **[[Programming:Registration]]** — the mechanics of the accounts a sock-puppet study needs.
design/platforms.1787862696.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki