6 — Data Sources and Source-Quality Framework
Phase 1 deliverable · research date 2026-09-15 · registry snapshot: 603 records
Part 1 — The source map
6.1 What the registry is
/home/claude/research/data/consolidated/sources.json holds 603 source records — the
sources a Phase 2 pipeline would actually poll, not a bibliography. Each was registered by
the sector analyst who used it, against the schema in §7 of the research contract, and each
carries 21 fields including reliability, bias, update frequency, access method, cost,
geographic coverage, historical depth, citation value, legal use restrictions and
monitoring priority.
| Measure | Value |
|---|---|
| Records | 603 |
| Sectors covered | 25 (22–25 records each; min 22 at sectors 02, 10, 13, 18, 19; max 25) |
| Distinct publishers | 494 |
| Distinct domains | 454 |
Distinct (name, publisher) pairs |
597 |
| Distinct source types | 24 of the schema's 27 enumerated values, plus one off-schema value (trade_org, 1 record, S-05-17) |
| Machine-readable (API or RSS) | 386 (64.0%) |
| Neither API nor RSS | 217 (36.0%) |
Non-null feed_url |
317 (52.6%) |
| Free at point of use | 497 (82.4%) |
| Requiring payment of any kind | 106 (17.6%) |
Deliberate duplication is a feature of the registry and a requirement on the pipeline.
603 records resolve to 454 domains because the same primary source was independently
registered by several sector analysts: www.sec.gov, data.sec.gov and efts.sec.gov
together account for 33 records across 22 of the 25 sectors; www.census.gov 14;
www.bls.gov 11; www.federalreserve.gov 8; www.federalregister.gov 7. Phase 2 must
deduplicate by endpoint for polling and fan out by sector for routing — one EDGAR
crawler, 22 sector subscriptions — or it will hit the SEC's 10-requests-per-second fair-access
limit 33 times over with 33 separate user agents.
6.2 Composition
By source type
| Type | n | % | Typical reliability | API | RSS | Default monitoring tier |
|---|---|---|---|---|---|---|
government_data |
130 | 21.6% | high (121/130) | 43% | 44% | mixed (45 t1 / 44 t2 / 41 t3) |
regulatory_filing |
92 | 15.3% | high (90/92) | 52% | 76% | tier1 (53 t1) |
trade_publication |
73 | 12.1% | medium (43/73) | 1% | 82% | tier2 (36 t2 / 24 t1) |
company_ir |
63 | 10.4% | high (59/63) | 3% | 48% | tier1–2 |
market_report |
58 | 9.6% | medium (34/58) | 14% | 29% | tier3 (33 t3) |
earnings |
30 | 5.0% | high (30/30) | 20% | 67% | tier2 |
survey |
27 | 4.5% | mixed (12 high / 15 med) | 11% | 41% | tier3 (19 t3) |
academic |
25 | 4.1% | high (24/25) | 20% | 76% | tier2–3 |
press_release |
24 | 4.0% | mixed (12/12) | 0% | 50% | tier1–2 |
public_dataset |
19 | 3.2% | high (12/19) | 37% | 11% | tier2 |
standards_body |
14 | 2.3% | high (13/14) | 7% | 50% | tier3 |
startup_db |
11 | 1.8% | medium (10/11) | 64% | 55% | tier2 |
pricing |
10 | 1.7% | high (8/10) | 30% | 10% | tier1 (7/10) |
trade_data |
6 | 1.0% | high (4/6) | 17% | 0% | tier1–2 |
newsletter |
4 | 0.7% | mixed | 0% | 100% | tier2 |
developer_activity |
3 | 0.5% | high | 100% | 33% | tier1–2 |
procurement |
3 | 0.5% | high (3/3) | 33% | 33% | tier1 |
app_store |
3 | 0.5% | medium | 67% | 0% | tier2 |
web_traffic |
2 | 0.3% | mixed | 50% | 50% | tier1–2 |
vc_announcement |
2 | 0.3% | medium | 50% | 100% | tier1–2 |
job_postings |
1 | 0.2% | medium | 0% | 0% | tier2 |
open_source |
1 | 0.2% | high | 0% | 0% | tier2 |
trade_org (off-schema) |
1 | 0.2% | medium | 0% | 0% | tier3 |
conference |
1 | 0.2% | low | 0% | 0% | tier3 |
patent |
0 | — | — | — | — | absent |
search_trends |
0 | — | — | — | — | absent |
social |
0 | — | — | — | — | absent |
community |
0 | — | — | — | — | absent |
The top four types are 358 records — 59.4% of the registry. The registry is overwhelmingly a primary-document registry: government statistics, regulatory filings, company disclosure and named-byline trade press. That is the correct shape for a platform whose evidence cap punishes weak sourcing, and it is a direct consequence of the tiering rules in §6.5. It is also a coverage bias that Phase 2 must correct deliberately (§6.4).
By reliability, cost, access and cadence
| Reliability | n | % | Cost | n | % | |
|---|---|---|---|---|---|---|
| high | 450 | 74.6% | free | 497 | 82.4% | |
| medium | 152 | 25.2% | freemium | 78 | 12.9% | |
| low | 1 | 0.2% | paid_high | 17 | 2.8% | |
| enterprise | 9 | 1.5% | ||||
| paid_low | 2 | 0.3% |
| Access method | n | % | Update frequency | n | % | Monitoring tier | n | % | ||
|---|---|---|---|---|---|---|---|---|---|---|
| public_page | 336 | 55.7% | quarterly | 130 | 21.6% | tier2_weekly | 246 | 40.8% | ||
| rss | 101 | 16.7% | daily | 124 | 20.6% | tier1_daily | 200 | 33.2% | ||
| api | 91 | 15.1% | irregular | 111 | 18.4% | tier3_monthly | 157 | 26.0% | ||
| bulk_download | 62 | 10.3% | monthly | 100 | 16.6% | |||||
| licensed | 12 | 2.0% | weekly | 69 | 11.4% | |||||
| 1 | 0.2% | annual | 48 | 8.0% | ||||||
| realtime | 20 | 3.3% | ||||||||
| semi-annual | 1 | 0.2% |
Cross-cuts that matter for engineering:
- 404 of 603 records (67%) are both
highreliability andfree. The intelligence product is overwhelmingly buildable on free primary sources. The licensing budget in §6.9 exists to close specific, named holes — not to buy the backbone. public_pageat 55.7% is the pipeline's central problem. More than half the registry must be fetched as HTML, which is exactly the population that the access-blocker register (§6.10) shows is most fragile.- 45 of the 200 tier-1 daily sources have neither an API nor an RSS feed. These are the highest-frequency, highest-value, least-automatable sources in the registry and they should be the first engineering allocation.
- Only 12 records are
licensed, and 10 of those aremarket_report. Licensing is concentrated, which is why §6.9 can be a short, prioritised list rather than a subscription sprawl.
By geography
| Coverage bucket | n | % |
|---|---|---|
| Global | 252 | 41.8% |
| US-only | 249 | 41.3% |
| EU / Europe | 40 | 6.6% |
| Asia | 29 | 4.8% |
| Other / regional | 33 | 5.5% |
41.3% US-only is the registry's largest structural bias, and the "global" bucket overstates coverage: 34 of the 252 global records carry an explicit skew in the field itself ("Global, US-weighted", "Global, US-skewed", "Global with strong Taiwan, Korea and China coverage"). The contract's non-English requirement did produce real regional depth — 48 records from non-US publishers including TSMC, SK hynix, Samsung, TrendForce, DigiTimes, Nikkei Asia, the Taiwan Stock Exchange MOPS (S-03-23), Korea Customs/MOTIE (S-03-24), CAAM (S-10-02), CPCA via CnEVPost (S-10-03), Gasgoo (S-10-14), China GACC (S-09-19), NPCI (S-07-23), VDMA (S-09-17), Eurostat (S-09-18), ENTSO-E (S-05-11), KOCCA (S-14-19), NPPA (S-15-06), JFTC (S-15-22), JARA (S-13-16), CRIA/MIIT (S-13-17), CONAB (S-19-20), the Federation of the Swiss Watch Industry (S-23-22) — but that is 8% of the registry against sectors whose centres of gravity are frequently outside the US.
6.3 Category-by-category evaluation
Each category below is assessed on the twelve attributes the brief requires. The compact table carries the quantitative attributes; the prose carries bias, data quality, historical depth, citation value and legal restrictions, which do not compress.
Summary table
| Category (registry type) | n | Reliability | Update | Access | Cost | API/RSS | Geo | Citation value |
|---|---|---|---|---|---|---|---|---|
Government publications (government_data) |
130 | high | monthly 50, daily 18, weekly 17 | public_page 54, bulk 33, api 30 | free 126 | 43% / 44% | US-heavy, strong EU/Asia tail | high |
Regulatory filings (regulatory_filing) |
92 | high | daily 29, irregular 32, realtime 13 | api 43, public_page 35 | free 91 | 52% / 76% | US-dominant | high |
Trade publications (trade_publication) |
73 | medium | daily 50 | rss 41, public_page 30 | free 45, freemium 21, paid 6 | 1% / 82% | global | medium |
Company sites & IR (company_ir) |
63 | high | quarterly 45 | public_page 51 | free 63 | 3% / 48% | global | high |
Market reports (market_report) |
58 | medium | quarterly 17, annual 13, monthly 12 | public_page 40, licensed 10 | free 29, paid/ent 15 | 14% / 29% | global | medium |
Earnings (earnings) |
30 | high | quarterly 30 | public_page 26 | free 30 | 20% / 67% | global | high |
Consumer/market surveys (survey) |
27 | mixed | annual 12, monthly 6 | public_page 22 | free 19 | 11% / 41% | US-heavy | medium |
Academic papers (academic) |
25 | high | weekly 6, irregular 7 | public_page 11, rss 7, bulk 5 | free 20 | 20% / 76% | global | high |
Press releases / product announcements (press_release) |
24 | mixed | irregular 14 | public_page 18, rss 6 | free 24 | 0% / 50% | global | medium |
Public datasets (public_dataset) |
19 | high | daily 8, quarterly 4 | bulk 8, public_page 9 | free 14 | 37% / 11% | global | high |
Standards bodies (standards_body) |
14 | high | irregular 9 | public_page 11 | free 13 | 7% / 50% | global/EU | high |
Startup databases (startup_db) |
11 | medium 10/11 | daily 5, quarterly 4 | api 4, public_page 5 | freemium 8, paid 2 | 64% / 55% | global, US-weighted | medium |
Pricing (pricing) |
10 | high | daily/weekly/monthly | public_page 8 | free 6, freemium 4 | 30% / 10% | global | high |
Import/export (trade_data) |
6 | high | monthly 6 | public_page 4 | free 4 | 17% / 0% | US, CN, CH, EU | high |
Newsletters (newsletter) |
4 | mixed | weekly 3 | rss 3 | freemium 3 | 0% / 100% | global | medium |
Developer activity (developer_activity) |
3 | high | daily 2 | api 2 | free 3 | 100% / 33% | global | medium |
Procurement (procurement) |
3 | high | daily 2 | public_page 2, api 1 | free 3 | 33% / 33% | US federal | high |
App stores (app_store) |
3 | medium | realtime/weekly/monthly | public_page 3 | free 1, freemium 2 | 67% / 0% | global | medium |
Web traffic (web_traffic) |
2 | mixed | realtime/monthly | api 1, public_page 1 | freemium 1, free 1 | 50% / 50% | global / US | medium |
VC announcements (vc_announcement) |
2 | medium | weekly 2 | rss 2 | freemium 2 | 50% / 100% | global | medium |
Job postings / hiring (job_postings) |
1 | medium | weekly | public_page | free | 0% / 0% | US | medium |
Open source (open_source) |
1 | high | irregular | public_page | free | 0% / 0% | global | medium |
Conferences (conference) |
1 | low | annual | public_page | paid_high | 0% / 0% | global | low |
| Patents | 0 | — | — | — | — | — | — | — |
| Search trends | 0 | — | — | — | — | — | — | — |
| Social | 0 | — | — | — | — | — | — | — |
| Community | 0 | — | — | — | — | — | — | — |
| M&A | 0 as a type | covered via regulatory_filing, earnings, trade_publication |
Government publications — 130 records, the registry's backbone
Named: EIA Short-Term Energy Outlook and Electric Power Monthly (S-05-01, S-05-02 — API + RSS, monthly, back to 1990/1997); US Census Value of Construction Put in Place C30 (S-09-01, S-18-01, bulk, back to 1993); Census Advance Monthly Retail Trade (S-12-02, S-19-09); Census Business Trends and Outlook Survey with AI questions (S-01-08, S-09-21, weekly, back to 2022); BLS CPI (S-19-01, back to 1913), PPI (S-18-05, back to 1947), Employment Situation (S-18-04, S-22-02, back to 1939/1948), JOLTS (S-22-04, back to 2000); Federal Reserve G.17 Industrial Production (S-09-03, back to 1919) and H.8 (S-07-02, back to 1973); FRED's Indeed software-development job postings index (S-02-04) and BVP Emerging Cloud Index (S-02-05); CISA advisories and the KEV catalog (S-04-01, S-04-02); FDA Novel Drug Approvals (S-06-01) and ClinicalTrials.gov (S-06-05, API, back to 1999); USAspending.gov (S-08-03, API, back to FY2008); ERCOT large-load interconnection queue (S-05-06, weekly, API + RSS); CAAM (S-10-02) and CPCA via CnEVPost (S-10-03); FAO Food Price Index (S-19-07, back to 1990); USDA WASDE (S-19-05, monthly since 1973); EPA Greenhouse Gas Reporting Program (S-20-11, bulk, back to 2010); IPEDS (S-22-05).
Reliability high (121 of 130). Bias: not commercial, but definitional and political —
the statistical agency decides the category, and the category can hide the story. Phase 1
confirmed two traps that any pipeline must encode: US Census reports data-centre
construction inside the "Office" category, which is why Census office construction was
+21.3% y/y in the middle of an office-distress narrative; and core CPI 2.4% (Aug 2026)
against core PCE ~3.3%, an unusually wide gap in the unusual direction. Never treat
"inflation" as one number; always name the index. Data quality high; revisions are
published and must be ingested as revisions rather than as new facts. Historical depth
is the category's decisive advantage — several series run 50–100 years, which is what makes
calibration_basis: base_rate possible at all (§9.7). Citation value high. Legal:
US federal works are public domain; the binding constraints are operational, not legal —
SEC requires a declared User-Agent with contact details and 10 req/s; most agencies expect
attribution. API/RSS: 43% / 44%, the best of any large category, and 30 records already
expose a real API.
Regulatory filings — 92 records, the highest-tier-1 concentration
Named: SEC EDGAR, registered 33 times across 22 sectors (S-01-01, S-07-04, S-13-03, S-21-01 and others) — full text back to 2001, filings back to 1993/94, plus the XBRL company-facts API (S-19-16, S-13-03); Federal Register with API (S-09-13, S-19-12, S-21-20, back to 1994 in full text); FERC eLibrary (S-05-04, daily, API + RSS, back to 1981); NRC ADAMS (S-05-13); NHTSA Standing General Order incident data (S-10-09); California DMV autonomous-vehicle disengagement reports (S-10-07) and CPUC quarterly AV reports (S-10-06); FCC IBFS satellite filings (S-08-14); ESMA MiCA interim register (S-21-08, weekly CSVs); BIS export-control actions (S-03-12); EUR-Lex Official Journal L series (S-20-01, back to 1952); European Parliament Legislative Observatory (S-20-02); Taiwan Stock Exchange MOPS (S-03-23 — mandatory monthly revenue filings, the best high-frequency semiconductor disclosure anywhere); CourtListener/RECAP (S-16-22, S-17-14 — free API, realtime).
Reliability high (90 of 92) — this is the Tier A core. Bias: none in the record;
substantial in what is absent. Confidential draft S-1 submissions do not appear until
public filing, so absence of a filing is not absence of an IPO process (S-01-01's own
bias note). Sealed filings are absent from RECAP by definition. Historical depth
excellent. Citation value the highest in the registry. Legal: public domain or
open-licence; rate limits and declared user agents are mandatory. API/RSS 52% / 76% —
the most machine-readable category. Access caveat: EDGAR's browse-edgar CGI is
robots-disallowed and efts.sec.gov 403s through some proxies; the working route is
data.sec.gov/submissions/CIK##########.json → fetch under /Archives/ (§6.10).
Company sites, IR, product announcements and press releases — 87 records combined
Named IR: NVIDIA (S-01-02, S-03-07), the hyperscaler set (S-01-03), TSMC with monthly revenue disclosure (S-03-01), SK hynix (S-03-03), Samsung (S-03-04), Micron (S-03-05), Broadcom (S-03-08), ASML (S-03-09), SpaceX (S-08-05 — first IR disclosure from Q2 2026), PJM Inside Lines (S-05-05), Netflix (S-14-03), Warner Bros. Discovery (S-14-07), Roblox (S-15-03, S-17-15), Meta (S-17-04), Snap (S-17-09), Reddit (S-17-20), Pinterest (S-17-16), Richemont (S-23-03), L'Oréal (S-23-04), Estée Lauder (S-23-05), Inditex (S-23-24), Kering (S-23-25), Spotify 6-K via EDGAR (S-24-08), Deezer (S-24-10). Named press-release/product-announcement sources: Anthropic newsroom (S-01-04), OpenAI index (S-01-05), TSMC Latest News (S-03-02), FDA Press Announcements (S-06-02), Apple Developer News and regional fee schedules (S-15-01, S-17-03), YouTube Official/Creator blog (S-17-10), TikTok Newsroom (S-17-23), USDA (S-19-10), Suno blog (S-24-17), IATA Press Room (S-25-01), aggregate media-sector issuer RSS (S-14-24).
Reliability high (59 of 63 IR records) for what is disclosed; bias is the issuer's
own framing, and vendor product announcements are marketing claims by definition and must
carry the label. Data quality high for audited figures, low for anything in a slide
footnote. Historical depth 10–30 years for most listed issuers. Legal: copyright to
the issuer; press material is normally citable with attribution; do not redistribute images
or full text. API/RSS 3% / 48% — and this is the category's operational weakness: IR
index pages are overwhelmingly JavaScript-rendered and fail to a plain fetch, while
individual press-release detail pages and PDF links fetch reliably. This was confirmed
independently in sectors 13, 16, 17, 19, 21, 22, 23 and 25. Route accordingly (§6.10).
Earnings — 30 records, quarterly, uniformly high reliability
All 30 are reliability: high and all 30 are quarterly — the only category in the
registry with perfect internal consistency. Named: Micron (S-03-05), NVIDIA (S-03-07),
Broadcom (S-03-08), ASML (S-03-09), GE Vernova orders (S-05-12), Tencent and NetEase
interim reports (S-15-13), The Trade Desk (S-16-20), Meta (S-17-04), Snap (S-17-09),
Pinterest (S-17-16), Reddit (S-17-20), NIKE (S-23-07), adidas (S-23-08), lululemon
(S-23-10), Spotify (S-24-08). Bias: non-GAAP presentation and segment definition are
chosen by the issuer; segment boundaries move. Citation value high; legal as for IR.
20% API / 67% RSS — the RSS route is the right default, with EDGAR as the authoritative
backstop.
Academic papers — 25 records
Named: arXiv quant-ph plus Nature/Science quantum error-correction coverage (S-03-21, API + RSS); Nature Medicine (S-06-07); Epoch AI data hub (S-01-09, API + bulk, models database back to 1950); LBNL Queued Up interconnection-queue dataset (S-05-08); Penn Wharton Budget Model effective tariff rates (S-09-12); Rhodium Group China research (S-10-19); Duke Nicholas Institute (S-05-24); Cloud Security Alliance (S-01-23); Future of Privacy Forum legislative trackers (S-17-18). Reliability high (24/25); bias — preprints are unreviewed and arXiv volume tracks incentives as much as progress; think-tank "academic" output carries funder positions and must be read as such. Update frequency irregular to weekly. Historical depth deep. Citation value high. Legal: licences vary — arXiv per-paper licences, many journals are closed-access; do not circumvent journal paywalls; use the preprint or the abstract. 20% API / 76% RSS.
Patents and patent databases — 0 records. The registry's single largest gap.
The schema enumerates patent as a source type. No sector registered one. This is a
material omission rather than a judgement that patents do not matter: §7 of the Phase 1
methodology weights leading indicators — hiring, patents, procurement, supply-chain shifts —
above lagging ones, and the patent leg is simply absent from the source layer. Patents were
referenced in dossier prose (ASML and export-licensing context in sector 03; ISO/TC 299 work
items in 13) but never as a monitorable feed.
Phase 2 must register, at minimum: USPTO PatentsView and the USPTO Open Data Portal (bulk + API, free, US, back to 1976, public domain); EPO Open Patent Services / Espacenet (API, freemium, global families, attribution and rate limits); WIPO PATENTSCOPE (public pages, global PCT); Google Patents BigQuery public dataset (bulk, free tier). Expected attributes: reliability high, update weekly, access api/bulk, cost free–freemium, geographic coverage global with well-known filing-jurisdiction bias, data quality high but noisy as a trend signal (filing volume tracks legal strategy and subsidy regimes as much as invention), historical depth decades, citation value high, legal restrictions minimal for US, attribution and rate limits for EPO.
Search trends — 0 records
Also enumerated in the schema and also unregistered. Partly defensible: the contract is
explicit that search interest is an input to velocity only and that velocity is 7.5%
of the composite, so a search-trends feed can never be load-bearing. But the absence means
the platform currently has no instrument for the fad diagnostic — attention up, adoption
flat — which is how §1.3 defines a fad. Phase 2 should register Google Trends (public page /
unofficial API, free, global, rolling, relative not absolute values, low data quality as
a level, useful only as a within-series delta) and Wikipedia Pageviews (official API, free,
bulk, back to 2015, genuinely absolute counts, high data quality, public-domain licence) —
and mark both citation_value: low so they cannot be laundered into evidence.
Social and community — 0 records as types; 12 adjacent
No social or community source is registered. The adjacent coverage is: the EU DSA
Transparency Database (S-17-24 — European Commission, daily bulk download, one common
schema across TikTok, Meta, Snap, Pinterest and X, tier1_daily, free, high reliability);
platform newsrooms as press releases (S-17-10 YouTube, S-17-23 TikTok); platform IR (Meta
S-17-04, Snap S-17-09, Reddit S-17-20, Pinterest S-17-16); Ofcom Online Safety enforcement
(S-17-01); the Australian eSafety Commissioner (S-17-05); Pew Research (S-17-08, S-22-15);
and one genuine community artefact registered as a public dataset —
videogamelayoffs.com (S-15-08, community-maintained, medium reliability, irregular).
This is the right shape and the wrong volume. Platform-reported regulatory data is Tier
A–adjacent; raw social volume is not evidence of anything under this contract. The DSA
database was flagged by sector 17 as "the best structured source in its sector" and is
still unexercised. Phase 2 should wire it first among social sources. Community forums
(subreddits, Discord, HN) should be registered only with citation_value: low and
reliability: low, used for discovery and never for support — and only via official APIs
under their terms, never by scraping authenticated surfaces.
Startup databases, VC announcements and M&A — 13 records + no dedicated M&A type
Named: Crunchbase, registered seven times across sectors 01, 02, 04, 07, 08, 11 and 22 (S-01-24, S-02-14, S-04-24, S-07-15, S-08-20, S-11-07, S-22-13 — API + RSS, freemium to paid_high, global, US-weighted); Carta Data Desk (S-11-03 — free, quarterly, high reliability, US-primarily, the only free public source for down-round rates, bridge rounds, stage-level round counts, AI vs non-AI valuation splits and fund-level DPI by vintage); Dealroom (S-11-17, Amsterdam, API); Tracxn (S-11-18, Bengaluru); CRETI proptech funding (S-18-19); CDR.fyi (S-20-06); Sightline Climate / CTVC (S-20-16, RSS, weekly).
Reliability medium for 10 of 11 startup_db records — the only category where medium is the norm, and correctly so. Bias: coverage is a function of who self-reports; Crunchbase and KPMG differ by $50.4bn on the same half-year of global VC ($510B vs $560.4B) and by 5,000+ vs 8,440 on Q2 deal count, a definitional difference in deal inclusion rather than an error. Carta's sample is self-selected: US-weighted, skewed to smaller and earlier-stage companies and sub-$100M funds. Data quality medium; historical depth ~2018 onward for Carta, mid-2000s for the commercial databases. Citation value medium — always name the tracker. Legal: redistribution of underlying datasets prohibited; free news tiers citable with attribution. API 64% / RSS 55% — the best-instrumented small category.
M&A has no dedicated source type. It is covered through regulatory_filing (8-K Item
2.01, S-4, HSR-driven disclosures — Agility Robotics' S-4, filed 2026-09-04, CIKs
0002074973 and 0001727116, is the worked example), earnings, DOJ Antitrust press releases
(S-14-15), state AG filings (S-14-16), and trade press. This is adequate for verification
and inadequate for discovery: Phase 2 should add an M&A-typed feed keyed on EDGAR form
types and antitrust dockets rather than buying a deal database.
Job postings and hiring — 1 record typed, ~6 functional
Named: the only job_postings record is S-02-19 (TechCrunch AI layoff tracker plus US
state WARN notices, medium reliability, weekly, US). Functionally the hiring signal is
carried by government data: FRED's Indeed software-development job postings index (S-02-04,
API, daily, baseline Feb 2020), BLS JOLTS (S-22-04, back to 2000), BLS Employment Situation
construction detail (S-18-04), BLS Occupational Outlook and Employment Projections (S-02-06),
and NACE first-destination surveys (S-22-25).
Given that hiring is one of the four leading indicators the methodology explicitly weights above media coverage, one typed record is under-provisioned. The structural finding that came out of hiring data in Phase 1 was among the strongest in the programme — US construction unemployment at a record-low 3.1% despite a soft market, with fewer than half of US metros adding construction jobs year over year, confirming data-centre crowd-out from the labour side. Phase 2 should register state WARN feeds directly (free, bulk, US, high reliability, legally clean) and, if a commercial postings panel is required, treat it as a licensing decision with an explicit bias note: job-posting panels measure advertised demand, which diverges from hiring in both directions.
App stores and product launches — 3 records
Named: Steam AI-content disclosure fields and store metadata (S-15-09 — Valve, free,
realtime, tier1_daily, high reliability, global PC); Appfigures Insights (S-17-11) and
Appfigures/Sensor Tower public releases (S-15-25), both freemium with partial API.
Bias: store-side metadata is issuer-declared (Steam's AI disclosure is self-reported);
third-party estimators model downloads and revenue from sampled panels and do not publish
methodology. Data quality medium for estimates, high for declared metadata. Historical
depth shallow. Citation value medium. Legal: store ToS prohibit bulk scraping —
use official APIs and published reports only. Critical context: app-store economics
fragmented by jurisdiction inside nine months — the flat 30% commission was dismantled
or altered across the US, EU, Japan, China, Korea, UK and Brazil; Apple's EU 5% CTC applies
from 2026-10-01 and China 25%/12% from 2026-03-15. Any app-economy model must now be
jurisdiction-specific, which makes the fee-schedule pages themselves (S-15-01, S-15-02) a
higher-value monitoring target than download estimates.
Developer activity and open source — 4 records
Named: Hugging Face Hub API and blog (S-01-18 — API + RSS, daily, tier1_daily, free,
high reliability); package-registry and developer telemetry across npm, PyPI and GitHub
(S-02-22 — API, daily, medium reliability, token-gated and rate-limited); the agentic-commerce
protocol documentation from Stripe, OpenAI and Google (S-12-20); OWASP GenAI Security Project
(S-04-15, the sole open_source record). The Model Context Protocol specification (S-02-07)
sits under standards_body and functions as the integration tripwire for enterprise software.
100% of developer_activity records expose an API — the highest rate in the registry.
Bias: registry download counts are heavily inflated by CI systems and mirrors and are
close to useless as adoption measures without de-botting; GitHub stars are a promotion
metric. Data quality high for counts, low for interpretation. Legal: GitHub API
requires a token and rate-limits; npm and PyPI publish open datasets. Citation value
medium — a real leading indicator, easily gamed, never sole support.
Newsletters — 4 records, all RSS
Named: SemiAnalysis (S-03-18 — freemium, weekly, strong China and Taiwan, subscriber-only
for the substantive work); GameDiscoverCo (S-15-10 — high reliability, weekly,
tier1_daily); GameDev Reports (S-15-18); aggregated law-firm regulatory advisories from
Arnold & Porter, Latham and peers (S-06-25 — high reliability, free, irregular).
100% RSS availability, which makes them cheap to ingest. Bias: single-author
newsletters carry a single analytical prior and frequently a commercial one; law-firm
advisories are marketing for a practice group but are unusually reliable on what a rule
says and when it takes effect — which is exactly the thing vendor marketing gets wrong.
Citation value medium; use as pointers to primaries, not as primaries.
Trade publications — 73 records, the discovery layer
Named, high value: Utility Dive (S-05-15), BioPharma Dive (S-06-13), Retail Dive (S-12-22), Construction Dive (S-18-20) — the Industry Dive stable, free, daily, RSS, named bylines; The Register (S-02-12); TechCrunch (S-02-13); The Robot Report (S-13-05); SpaceNews (S-08-18); Breaking Defense (S-08-19); AdExchanger (S-16-11); Digiday (S-16-12); Data Center Frontier (S-18-21); DatacenterDynamics (S-01-12); PocketGamer.biz (S-15-16); Game Developer (S-15-17, back to 1994 as Gamasutra); CnEVPost (S-10-21) and Gasgoo (S-10-14) for China; South China Morning Post Tech (S-01-17); Nikkei Asia (S-03-17, metered); DigiTimes (S-03-15, subscription — and the likely origin of much uncredited Taiwan supply-chain reporting); TrendForce (S-03-14); STAT News (S-06-14) and Endpoints (S-06-15), both paywalled; Reuters Health and Pharmaceuticals (S-06-23, enterprise); Courthouse News antitrust (S-16-21).
Reliability medium for 43 of 73 — correctly graded down. 82% RSS availability makes this the cheapest high-volume ingest in the registry, and the reason trade press is the discovery layer rather than the evidence layer. Bias: advertiser and event revenue from the industry covered; Industry Dive, Informa and Access Intelligence all run conferences for the sectors they report on. Legal: copyright; headlines and short quotation with attribution only; no full-text redistribution and no paywall circumvention. Syndication is the central risk and is handled in §6.7.
Conferences — 1 record, and it is the registry's only low reliability source
S-16-24 (Cannes Lions / Advertising Week / Possible programmes and disclosures) is the sole
conference record, reliability: low, citation_value: low, cost: paid_high, and its
bias note is the correct general rule for the whole category: "Pay-to-present environments.
Almost every quantitative claim made on stage is vendor-sponsored and unaudited. Useful as a
leading indicator of narrative, explicitly not of adoption." Conference programmes are a
legitimate narrative instrument — they show what a sector has decided to talk about six
months ahead — and must never support a factual claim.
Consumer and industry surveys — 27 records
Named: Census Business Trends and Outlook Survey (S-01-08, S-09-21, S-13-14 — weekly, API, the only official US AI-adoption series); ISM Manufacturing PMI (S-09-07, back to 1948); Verizon DBIR (S-04-06); Stack Overflow Developer Survey (S-02-08); Pew Research (S-17-08, S-22-15); NAM Manufacturers' Outlook (S-09-16); Modern Materials Handling / Peerless intralogistics robotics surveys (S-09-24, S-13-19); WFA (S-16-15); CDP disclosure and scores (S-20-18); NACE (S-22-25); GBTA Business Travel Index (S-25-17, paywalled data cube).
Reliability is genuinely mixed (12 high / 15 medium) and this is the category where the
fact vs estimate line matters most. Under §1.3, survey-reported intent is an estimate;
transaction, traffic and usage data are fact — and the two diverge constantly, with
travel the worst offender. Bias: trade-association surveys have a membership interest in
the narrative; panel composition is rarely disclosed; year-over-year change within the same
panel is the usable signal, levels are not. The canonical trap from sector 25: GBTA's 2026
release shows spend +7.2% against trips +1.3%; quoting only the first is the sector's
most common analytical error. Legal: press releases quotable; data cubes licensed.
Market reports — 58 records, the most problematic category
Named, credible: SEMI WWSEMS (S-03-10), WSTS (S-03-11, back to 1986), IFR World Robotics (S-13-01 free releases / S-13-02 licensed reports), Nielsen The Gauge (S-14-01), Circana (S-15-11), BloombergNEF (S-05-19, S-10-13), WARC (S-16-17), EMARKETER (S-12-15), Comscore (S-14-12), Trepp (S-07-19, S-18-11), Green Street CPPI (S-18-12), PitchBook (S-07-21), Coresight (S-12-17), SNE Research (S-10-12), Madison & Wall (S-16-18), Grid Strategies (S-05-23), Tax Foundation tariff tracker (S-09-25), Marsh Global Insurance Market Index (S-04-09), KPMG Venture Pulse (S-07-16), Chainalysis (S-04-08).
Reliability medium for 34 of 58, tier3_monthly for 33 of 58, and 15 of the 28 strictly-paid
records in the whole registry (paid_low + paid_high + enterprise) sit here. The category is bimodal: a small set of
methodologically serious industry-body series (SEMI, WSTS, IFR, WSTS-derived), and a large
set of commercial forecasters whose free tier is marketing for the paid product.
The structural pattern to encode: headline free, detail paid. SEMI publishes the
$40.53bn global equipment total and withholds the regional table where China's trajectory is
visible. IFR publishes installations, stock and density and withholds end-use, application,
vendor-level and service-robot data. Coresight publishes a weekly summary and withholds the
line items. Green Street publishes monthly percentage changes and withholds index levels.
This is not an accident of pricing; it is the product design, and it means the free tier
is systematically sufficient for direction and systematically insufficient for level.
Trend detection can run on the free tier. Any claim about a level needs the licence or must
be labelled estimate with the modeller named.
Bias: commercial interest in the market looking large; methodology usually undisclosed; forecasts revised silently. Legal: redistribution prohibited almost universally; attribution required for headlines. Citation value medium, and Tier C on sight for anything without a named methodology (§6.6).
Public datasets — 19 records, the highest-leverage underused category
Named: EU DSA Transparency Database (S-17-24 — daily bulk, free, high reliability, common schema across all designated platforms); IMF PortWatch (S-09-23 — daily bulk, API, free, ~1,400 ports and major chokepoints, AIS-derived); CVE Program cvelistV5 (S-04-22 — use the git repository, not the website; back to 1999); Berkeley Voluntary Registry Offsets Database (S-20-09 — quarterly bulk, free, global, the working substitute for the JS-rendered Verra registry); Global Energy Monitor coal plant tracker (S-05-18); SIPRI arms industry and military expenditure databases (S-08-21, annual bulk); DefiLlama (S-21-10, realtime API), RWA.xyz (S-21-09), Artemis (S-21-11); Visa Onchain Analytics (S-21-12 — free, daily, high reliability, the only published methodology that separates genuine stablecoin payments from bot, bridge, MEV and exchange-internal traffic); IPEDS (S-22-05); National Student Clearinghouse enrolment (S-22-06); ITRC data-breach reports (S-04-07); NOAA Billion-Dollar Disasters (S-20-23); NCREIF (S-18-13, licensed); Box Office Mojo (S-14-10) and The Numbers (S-14-11).
37% expose an API; 8 of 19 are bulk downloads. Bias varies sharply: NOAA's series is authoritative and has stopped; crypto dashboards are self-published by interested parties and, per sector 21, "the sector's headline metrics are manufacturable and demonstrably manufactured" — an intelligence pipeline that ingests aggregator dashboards uncritically will import fabricated data at scale. Visa's dashboard is the counter-example precisely because its adjustment criteria are disclosed (exclude addresses exceeding 1,000 transactions or $10m in 30 days, CEX flows, mint/burn and infrastructure operations). Legal: mostly open licences with attribution. This category contains the registry's best cost-to-value ratio and its least-exercised assets.
Standards bodies — 14 records
Named: NIST CSRC post-quantum migration (S-04-17), NERC Reliability Assessments (S-05-03),
ENTSO-E Transparency Platform (S-05-11 — realtime API, 39 TSOs across 36 countries), Model
Context Protocol (S-02-07), FSB (S-07-09), ECB digital euro rulebook (S-07-12), ILPA
(S-11-19), ISO/TC 299 Robotics (S-13-20, paid_low — ISO sells its standards), Media Rating
Council (S-16-10), EFRAG ESRS (S-20-03), ICVCM (S-20-07), Verra (S-20-21), 1EdTech (S-22-24),
European Commission AI Act GPAI Code of Practice (S-24-24).
Reliability high (13/14); update irregular (9/14). Value: standards bodies are the
earliest formal signal of a regulatory trend, and they publish work items before rules
exist. Legal: ISO and several others sell the standard text — budget paid_low per
standard, and never republish clause text.
Procurement — 3 records, all US federal, all high reliability
DoD daily contract announcements over $7.5m (S-08-01 — RSS, daily, tier1_daily), SAM.gov
contract opportunities and awards (S-08-02 — API, daily), Space Force / SDA acquisition
announcements (S-08-15). Plus USAspending.gov under government data (S-08-03, API, back to
FY2008). Bias: US-only, and classified programmes are absent by construction — absence
of an award is not absence of a programme. Citation value high; legal: public domain.
Procurement is a named leading indicator in the methodology and is currently a single-country
instrument; EU TED and UK Contracts Finder are the obvious Phase 2 additions.
Import/export — 6 typed records plus substantial government-data coverage
Named: US FT-900 (S-09-04 — Census/BEA, API, monthly, full country and HS-code detail);
BLS import/export price indexes (S-09-05); China GACC monthly trade statistics (S-09-19);
China customs rare-earth and permanent-magnet export data mirrored by Silverado Policy
Accelerator (S-13-18 — monthly, tier1_daily, high reliability, by destination country);
Korea Customs / MOTIE semiconductor exports (S-03-24); OTEXA US textile and apparel trade
(S-23-23 — bulk, free, tier1_daily); Federation of the Swiss Watch Industry exports
(S-23-22); USITC HTS and CBP CSMS tariff guidance (S-03-13); Port of LA / Long Beach
container statistics (S-09-11).
This category punched far above its weight in Phase 1. Korean monthly export values are described in the sector 03 dossier as "historically the earliest reliable turn signal in this sector." Chinese rare-earth licensing data produced the finding that Japan received zero covered rare-earth exports in July 2026, which is a robotics and industrial constraint, not only a defence one. Bias: national statistics reflect national framing; transshipment makes origin attribution unreliable, and the split between genuine relocation and transshipment in Vietnam and Mexico flows was explicitly recorded as unresolvable from available data. Legal: government data, free. 0% RSS, 17% API — this category needs bulk-download plumbing, not feeds.
Pricing — 10 records, 7 of them tier1_daily
Named: Drewry World Container Index (S-09-08, weekly), Cass Transportation Index
(S-09-09, monthly), DAT spot rates and load-to-truck (S-09-10, weekly, partial API), World
Bank Pink Sheet commodity prices (S-19-08, monthly bulk + API, global), IATA Jet Fuel Price
Monitor (S-25-03, weekly), Kelley Blue Book average transaction price (S-10-04, monthly),
OpenRouter LLM rankings (S-01-11, API, daily), software vendor pricing and packaging pages
(S-02-16, tier1_daily — pricing-page diffing is the instrument for the per-seat-to-consumption
shift, T-02-01), Apple and Google fee schedules (S-15-01, S-15-02), BVP cloud index multiples
(S-02-21).
Highest tier-1 density of any category (7 of 10) because prices move daily and are the least-laggy observable in the registry. Bias: index construction is proprietary for Drewry, Cass and DAT; vendor list prices diverge from realised prices. Data quality high; historical depth decades for the commodity and freight series. Legal: Drewry restricts automated access via robots.txt and its full dataset is a licensing cost; World Bank is open. Gap: there is no audited CPM index for advertising in any channel, and no current independent price-per-token index — Epoch AI's inference-price dataset was last updated 2025-03-12 and Artificial Analysis's pricing table returned HTTP 429 repeatedly.
Web traffic — 2 records
Cloudflare Radar (S-04-18 — API, realtime, freemium, global with country breakdowns, high
reliability) and Adobe Digital Insights retail and AI-traffic reports (S-12-16 — monthly,
free, US, medium reliability, published by an interested vendor). This is thin for a
category that matters, particularly as AI-referred traffic becomes a live commercial
question in sectors 12, 16, 17 and 24. Bias: every commercial traffic estimator models
from a panel and none publishes the panel. Phase 2 should add Cloudflare Radar's full API
surface and the Wikimedia Pageviews API (genuinely counted, not modelled) and should mark
all panel-derived traffic estimates estimate, never fact.
6.4 The registry's own biases — read this before trusting the map
- US-centricity. 41.3% US-only; the "global" bucket contains 34 records with an explicit skew in the field itself. Sectors whose centre of gravity is outside the US (semis, EV/battery, luxury, K-content, shipping) met the regional-source requirement, but sector 16 recorded "China, India, Brazil and Southeast Asia: entirely absent" and sector 11 consulted no Chinese-language primary sources at all.
- Reliability grade inflation. 450
high, 152medium, 1low. A registry in which 0.2% of sources are low-reliability is not describing the information environment; it is describing a filtered environment. The filtering was correct — content farms were excluded rather than registered — but Phase 2 must not read the distribution as evidence that most sources are good. - Selection toward the fetchable. 82.4% free, 55.7% public-page. Sources that blocked retrieval are under-registered relative to their value; several were registered only so the gap would be auditable (S-20-25 CARB, "Tier A source, zero coverage").
- Discovery-narrow sectors. Sectors 13–18 completed with 5–11 searches each instead of ~22 because the shared search budget was exhausted. Their Tier-A verification is sound; their source discovery is narrow, and their registry entries should be treated as a floor.
- Four missing types (
patent,search_trends,social,community) and four near-empty ones (job_postings1,open_source1,conference1,web_traffic2) against a methodology that explicitly weights hiring and patents as leading indicators. - One off-schema value (
trade_org) indicates the enum is not validated on write. Fix at ingest.
Part 2 — The source-quality framework
6.5 A / B / C tiering and how to apply it
Tier A — primary. The record itself: SEC and other regulatory filings, company IR and press pages, government statistics agencies, central banks, standards bodies, patent offices, peer-reviewed papers, court documents. Tier B — reputable secondary. FT, WSJ, Reuters, Bloomberg, Nikkei, The Information, STAT News, trade press with named bylines, Big-4 and top-tier consultancy research. Tier C — everything else. Aggregators, vendor-sponsored market reports, SEO content farms, press-release wires, LLM-written blogs.
Three rules, all enforced rather than encouraged:
- Prefer A > B > C, always.
- Tier C may never be the sole support for a claim.
- Two outlets syndicating the same wire story are ONE source, not two.
Default tier by registry type
| Registry type | Default tier | Promotion / demotion |
|---|---|---|
regulatory_filing, government_data, standards_body, academic (peer-reviewed), company_ir, earnings, procurement, trade_data |
A | Demote to B if the record is a summary page rather than the document; demote if undated |
press_release |
A for facts about the issuer, marketing for claims about the market |
A vendor's claim about its own product is marketing and must carry the label |
public_dataset |
A if the methodology is published (Visa Onchain, Berkeley Offsets, CVE); B otherwise | Demote self-published crypto dashboards without disclosed adjustment criteria |
trade_publication |
B with a named byline; C without | Promote to A when it reproduces a primary document in full |
newsletter |
B if the author is named and the analysis is original | C if it aggregates without attribution |
market_report |
B if the modeller and methodology are named and dated; C otherwise | Industry-body statistical series (SEMI, WSTS, IFR) are B, and A for their own membership data |
survey |
B, claims recorded as estimate |
C if panel size and composition are undisclosed |
startup_db |
B, claims recorded as estimate |
Always name the tracker |
app_store, web_traffic, developer_activity, search_trends |
B for declared/counted data, C for modelled estimates | Never sole support |
conference |
C | Narrative signal only |
Applying the tier at ingest
- Every source needs a publication date. Undated web pages are Tier C at best and
carry
undated: true. Sector 01 found this is also a retrieval trap: "searching a current metric often surfaces a two-year-old article in the top results" — check the date of every result before use, and prefer the primary issuer's own current page. - Tier determines the evidence cap, and the cap is arithmetic.
evidence_quality: 5requires multiple Tier-A primaries; Tier-C-only caps it at 0–1;score = min(raw, 40 + 60E)whereE = mean(evidence_quality, source_diversity)/5. A trend with weak evidence cannot exceed 40 no matter how loud it is. 25 of the 500 seed trends are capped by this rule. source_diversitycounts independent organisations, not links: 1 org = 1, 2 = 2, 3–4 = 3, 5–6 = 4, 7+ across at least two source types = 5.
6.6 Content-farm detection
Content farms dominate the search results for every sector in this programme. Sector 13 recorded them as "worse here, because humanoids are the most SEO-contested robotics topic"; sector 02 recorded that the entire public space for SaaS multiples and NRR benchmarks is content-farm-controlled; sector 09 found reshoring, warehouse-automation and industrial-AI-ROI statistics to be "almost entirely undated, unauthored pages citing unsourced percentages."
The signature — any three of these is a quarantine
| Signal | Test |
|---|---|
| Title template | `/(Market (Report |
| No named author | No author meta, no byline, or a byline that is a brand ("SEO DIGITAL PROS") |
| No date, or a rolling date | Missing datePublished, or a date that changes on re-fetch |
| Unsourced forward figure | A $XX billion by 20NN claim with no named methodology, sample or modeller |
| Monoculture | The domain's entire content is market-size CAGR posts |
| Citation loop | The only corroboration is other domains matching the same signature |
| Wire-press origin | openpr.com and equivalents, which publish anything paid for |
| Statistic with no denominator | "89% of AI pilots fail" attributed to a consultancy with no retrievable primary |
Named and excluded in Phase 1
Recorded so the blocklist is auditable and seeded rather than rediscovered. From sector 06: lifesciencedaily.news, visionlifesciences.com, intuitionlabs.ai, peptidejournal.org, aimagicx.com, pdpspectra.com, humai.blog, presenc.ai, aimmediahouse.com, corstrate.com, nextaipress.com, deepceutix.com, biotechsign.com, hcranking.com, rxalmanac.com — several of which carried the only claims found on Isomorphic Labs' 2026 funding and on AI-drug-discovery market sizing, and were not laundered into the dossier. From sector 01: the "89% of enterprise agent pilots never reach deployment" figure, whose clearest attributing article carries the byline "SEO DIGITAL PROS" and no link, sample size or methodology — cited in the dossier only as an example of unsourced statistic propagation. From sector 11: search results for startup failure rates "dominated by Tier-C content farms recycling a decades-old '90% fail' claim."
Operational specification
- Deny-list the named domains and the
openpr.comclass. Blocking, not down-ranking. - Score every new domain on the eight signals at first ingest; ≥3 →
tier: C,quarantine: true. - Quarantine is not deletion. Quarantined items remain retrievable as evidence of what is circulating, which is itself a signal — the divergence between what is claimed and what is verifiable is the fad diagnostic.
- Hard rule: a quarantined source can never be the sole support for a claim, and can
never raise
evidence_qualityabove 1. - Track propagation. When a figure appears across ≥5 quarantined domains within 30 days with no primary, flag it as a propagating unsourced statistic and publish it as such. This is a product feature, not just a filter.
6.7 What independence actually means
source_diversity counts independent organisations. Four failure modes destroy apparent
independence, and all four were observed in Phase 1:
- Wire syndication. The same story appears across dozens of domains. Sector 17 states the rule: "Syndication means the same story appears across many domains, which can create a false impression of independent corroboration — under this contract that is one source, not several."
- Shared upstream data. PitchBook underpins KPMG's venture figures (S-07-21), so citing both is not triangulation. EMARKETER and WARC observe the same advertiser panel universe (S-12-15) and must not be treated as independent confirmation of each other. Endpoints News is cited by other trackers as an underlying data source (S-06-15). Madison & Wall's concentration estimates were the most widely syndicated in 2026, so "apparent multi-source agreement often traces back to this one shop" (S-16-18).
- Republication chains. CnEVPost republishes both CPCA data and SNE Research battery tables; Silverado republishes Chinese customs data; the ISM PMI is most reliably obtained through its PR Newswire distribution. These are useful — often more machine-readable than the primary — but they are the primary, not a second source.
- Self-referential estimates. A vendor's market-size figure quoted back by the trade press that the vendor advertises in.
Dedupe algorithm for Phase 2:
for each incoming claim:
normalise the numeric assertion (figure, unit, period, subject)
cluster by (assertion, ±48h)
within the cluster:
resolve origin = earliest publication with an original byline,
or the named wire/agency, or the named data vendor
all other members -> republication, weight 0 for source_diversity
members citing a shared upstream vendor -> collapse to that vendor
source_diversity = count(distinct origin organisations)
across >= 2 distinct source_types for a score of 5
Record the republication chain rather than discarding it — chain length and speed is a
useful measure of narrative velocity, and is exactly the input velocity is allowed to take.
6.8 Paywalled and licensed data
The rule, restated without qualification:
No scraping behind logins. No bypassing access controls. No circumvention of paywalls. No ToS violations. Prefer official APIs, RSS feeds, public pages and licensed data.
This is binding on Phase 2 engineering as it was on Phase 1 research, and it was observed: sector 01 read only the headline and dek of The Information's Anthropic gross-margin story and recorded the body as unread rather than obtaining it; sector 06 recorded STAT+ methodology as unverifiable rather than extracting it. An honest gap beats a plausible fabrication, and it also beats a ToS breach.
Four legitimate routes, in preference order:
- Free tier with attribution. Sufficient for direction in most cases. The free IFR press releases carry global installs, stock and density; the free SEMI release carries the global total; Green Street publishes monthly percentage changes; Circana headlines are released publicly via ESA and named public commentary.
- Licence the data. §6.9.
- Substitute. Berkeley's Voluntary Registry Offsets Database (S-20-09, free bulk) in place of the JS-rendered Verra registry; RBI payment-system indicators in place of NPCI; FFIEC CDR bulk in place of the FDIC QBP web pages; Justia (S-24-06) where CourtListener is blocked; on-chain supply in place of JS-rendered issuer dashboards.
- Record the gap. Publish the absence as a finding. Phase 1's most valuable negative results are of this kind: no frontier AI lab discloses a gross margin; no operator anywhere publishes robotaxi unit economics; no independent audited DPI benchmark exists for venture capital; no audited CPM index exists for any advertising channel.
6.9 Prioritised Phase 2 licensing budget
Ranked by: (a) no free substitute exists for the load-bearing series, (b) number of sectors
served, (c) whether a scored trend currently depends on it, (d) monitoring cadence. List
prices are not published for most of these and must be quoted; the cost band from the
registry is given instead.
Tier 1 — buy first. No substitute; a scored trend depends on it.
| # | Source | ID | Band | Sectors | What the licence buys that free does not |
|---|---|---|---|---|---|
| 1 | IFR World Robotics (Industrial + Service) | S-13-02 | paid_high | 13, 09 | End-use, application and vendor-level breakdowns; the full national time series; the entire service-robot dataset. Free releases give only installs, stock and density. Annual series back to 1993. "Budget for this licence; there is no substitute." Sector 13 scores intelligence_demand 5 against data_availability 3 — maximum demand, worst information environment. |
| 2 | SEMI WWSEMS regional data | S-03-10 | freemium→paid | 03, 09 | The seven-region equipment table — the part that reveals China's equipment purchasing trajectory. The free release gives only the global total ($40.53bn, Q2 2026). Back to the 1990s. |
| 3 | WSTS detailed billings | S-03-11 | freemium→member | 03 | Monthly billings by product category and region, back to 1986. Member-restricted. "This is precisely why content-farm 'market size' pages fill the search results" — buying WSTS is buying the ability to refuse Tier C. |
| 4 | PitchBook / Morningstar | S-07-21 | enterprise | 07, 11, 02, 18, 22 | The private-markets database behind the $607bn / 567-fund evergreen universe. Note: it also underpins KPMG's venture figures — licensing it does not add an independent source, it identifies the shared one. |
| 5 | WARC | S-16-17 | enterprise | 16, 12 | Global ad-spend methodology and the effectiveness case database. The only source that surfaced independent academic evidence on AI creative performance rather than vendor lift claims. Sector 16's own instruction: "licensing budget should go to WARC first." |
| 6 | STAT+ (STAT News) | S-06-14 | paid_high | 06 | The M&A running totals used in sector 06 ($134bn YTD to 2026-06-22) sit behind it with their methodology. Per-seat licence; do not circumvent. |
| 7 | Circana US video game consumer spending | S-15-11 | paid_high | 15 | The authoritative US spending series back to the 1990s under NPD. Only headlines are free, via ESA. |
| 8 | Coresight store openings/closures databank | S-12-17 | paid_high | 12, 18 | The only weekly cadence in retail footprint data; line-item detail behind the paywall (the public preview gives ~7,900 closures / 5,500 openings for full-year 2026). |
Tier 2 — conditional. Buy if the theme is tracked.
| # | Source | ID | Band | Sectors | Trigger for purchase |
|---|---|---|---|---|---|
| 9 | BloombergNEF (Energy Transition Investment Trends, NEO, battery price survey) | S-05-19, S-10-13 | paid_high | 05, 20, 10 | If energy-transition investment or battery cost curves are tracked. BNEF's data-centre demand trajectory (1,114 TWh by 2050) is materially more conservative than US utility forecasts — that divergence is itself worth tracking |
| 10 | Trepp CMBS loan-level | S-07-19, S-18-11 | enterprise | 07, 18 | If CRE credit is a tracked theme. Headline delinquency rates are free; loan-level and history are not |
| 11 | MSCI climate and carbon markets | S-20-24 | enterprise | 20 | If carbon prices or VCM volumes are needed. MSCI sells the weekly carbon credit price index and now gates First Street's research library. Together with Sylvera, BeZero and Abatable this is where VCM price and volume data actually lives, and none of it is free. Its own free pages are stale marketing (Net-Zero Tracker still dated 2024-08-31) |
| 12 | Comscore Box Office | S-14-12 | enterprise | 14 | If theatrical is tracked. The industry's actual currency; the public page returned no usable figures at all |
| 13 | The Information | S-01-19 | paid_high | 01, 02 | Currently the only outlet reporting frontier-lab gross-margin projections and cash-flow-positive timing |
| 14 | Green Street CPPI + NCREIF | S-18-12, S-18-13 | enterprise / paid_high | 18, 07 | Buy as a pair with a transacted index: REIT-implied (Green Street), appraised (NCREIF) and transacted. The spread between the three is more informative than any one of them |
| 15 | EMARKETER | S-12-15 | paid_high | 12, 16 | Only with WARC, and only with the independence caveat recorded on the claim |
| 16 | Crunchbase (data, not news) | S-02-14 | paid_high | 01, 02, 04, 07, 08, 11, 22 | Serves seven sectors; but Carta (free) covers the benchmark questions better and the Crunchbase/KPMG divergence must be carried regardless |
| 17 | DigiTimes / Nikkei Asia | S-03-15 / S-03-17 | paid_high / paid_low | 03, 09 | DigiTimes is the likely origin of much uncredited Taiwan supply-chain reporting — licensing it collapses several apparent sources into one, which is a quality purchase, not only an access one |
| 18 | Endpoints News, Secondaries Investor, GBTA BTI Data Cube, ConstructConnect, SNE Research, NCREIF, Aviation Week AWIN | S-06-15, S-11-20, S-25-17, S-18-22, S-10-12 | paid_high / enterprise | 06, 11, 25, 18, 10, 08 | Sector-specific; each named in its dossier's data-gaps section. SNE's free republished headlines are usually sufficient for detection |
Tier 3 — defer
ISO/TC 299 standards text (S-13-20, paid_low, buy per standard on demand); conference
programmes and access (S-16-24 — reliability: low, citation_value: low, buy only as a
narrative instrument and never as evidence).
Not a licensing problem — an engineering problem
Four sources the brief and the dossiers flag as unobtainable are free, and were missed for retrieval reasons. Spending licence money on them would be an error:
| Source | ID | Status | Actual fix |
|---|---|---|---|
| Visa Onchain Analytics adjusted stablecoin volume | S-21-12 | free, daily, high reliability | The dashboard is JS-rendered and returned methodology text but no figures. Fix: headless browser, or the underlying Allium data. This is the only source with a published methodology for separating genuine stablecoin payments from bot, bridge, MEV and exchange-internal traffic |
| CourtListener / RECAP | S-16-22, S-17-14 | free, official API, realtime | Robots-disallowed through the Phase 1 session proxy. Fix: API token + declared user agent + rate limiting, per Free Law Project terms. Flagged in sector 16 as "HIGHEST-PRIORITY ADDITION FOR PHASE 2" — the Google remedies memorandum and proposed final judgment land on this docket, not in any press release |
| Carta Data Desk | S-11-03 | free, quarterly, high reliability | Not paywalled at all. The gap is a publication gap — Carta had not published a Q2 2026 State of Private Markets at research date. Fix: monitor, don't buy. It remains the only free public source for down-round rates, stage-level round counts, AI vs non-AI valuation splits and fund-level DPI by vintage |
| CISA KEV / CVE Program | S-04-02, S-04-22 | free, machine-readable by design | KEV JSON returned 403 to generic HTTP clients through some proxies; fix with a browser user-agent. For CVE use the cvelistV5 git repository, not the website |
Highest-value free datasets still unexercised, which should be wired before any licence is signed: EU DSA Transparency Database (S-17-24), ESMA MiCA interim register (S-21-08), IMF PortWatch (S-09-23), OTEXA (S-23-23), SEC XBRL company-facts API (S-13-03, S-19-16), Berkeley Voluntary Registry Offsets Database (S-20-09), EIA API (S-05-01/S-20-14), Eurostat education and tourism (S-22-17, S-25-10), ENTSO-E Transparency Platform (S-05-11).
6.10 Access-blocker register
Engineering guidance. Every entry is a source the research needed, could not retrieve, and has a legal route to. None of the recommended routes involves circumventing an access control — every one is an alternative published endpoint, an official API, or a licence.
Class A — robots-disallowed
| Source | ID | Blocker | Recommended route |
|---|---|---|---|
SEC EDGAR browse-edgar |
S-01-01, S-07-04, S-13-03, S-14-08, S-16-08, S-17-02, S-21-01, S-22-12, S-23-21, S-24-07, S-25-05 | /cgi-bin/browse-edgar disallowed; efts.sec.gov 403s through proxies |
data.sec.gov/submissions/CIK##########.json to list filings → fetch the document under /Archives/. Use the XBRL company-facts API for figures. Declared User-Agent with contact, ≤10 req/s |
| congress.gov / CRS reports | S-08-12 | robots-disallowed | govinfo.gov bulk data; Federal Register API for enacted instruments |
| NPCI UPI statistics | S-07-23 | npci.org.in robots-disallowed |
RBI monthly payment-system indicators or an official data partner |
| ISM Manufacturing PMI | S-09-07 | ismworld.org robots-restricted |
The PR Newswire distribution of the same release — machine-readable |
| Drewry World Container Index | S-09-08 | robots-restricted | Licence, or secondary republication (Daily Cargo News). Fragile — treat as a licensing candidate |
| California Air Resources Board | S-20-25 | ww2.arb.ca.gov robots-disallowed; Tier A source, zero coverage |
California Regulatory Notice Register; CARB board-meeting PDFs by direct path; EPA/Federal Register for federal analogues |
| EUR-Lex search | S-20-01 | eur-lex.europa.eu/search.html robots-disallowed |
Direct legal-content URLs; find act numbers first via the EP Legislative Observatory (S-20-02) |
| OCC | S-07-07, S-21-21 | every occ.gov/occ.treas.gov path ROBOTS_DISALLOWED; robots.txt itself timed out |
Federal Register API for OCC rulemakings and bulletins; FFIEC for supervisory data |
| European Commission digital-strategy sub-paths | S-17-06 | several sub-paths robots-disallowed | EUR-Lex + the Commission press-corner RSS |
| eSafety Commissioner (AU) | S-17-05 | robots.txt fetch timeout | Direct PDF paths and the publications RSS |
| Swiss watch industry exports | S-23-22 | fhs.swiss robots failures |
Swiss Federal Office for Customs (Swiss-Impex) |
| OTEXA | S-23-23 | otexa.trade.gov robots connect timeout |
Bulk download from an unblocked egress; this is the single highest-value unexercised dataset in sector 23 |
| Inditex, Kering | S-23-24, S-23-25 | hard site blocks; the sector's largest fast-fashion company contributed no 2026 data | CNMV (Spain) and AMF / Euronext regulated-information filings — the statutory disclosure, not the corporate site |
| SpaceNews | S-08-18 | robots restricts some paths | RSS feed only |
| CourtListener | S-24-06 note | robots-disallowed through the session proxy | API token + declared UA (see §6.9); Justia (S-24-06) as interim substitute |
Class B — JavaScript-rendered
| Source | ID | Blocker | Recommended route |
|---|---|---|---|
| Corporate IR index pages, sector-wide | S-15-03, S-16-04, S-16-07, S-17-15, S-17-16, S-17-20, S-19-17, S-22-23, S-23-01, S-25-08, S-25-24, S-13-07, S-13-08, S-11-14, S-16-20 | Index pages render client-side and return navigation shells | Go straight to the individual press-release detail page or PDF — these fetch reliably. Confirmed independently in sectors 13, 16, 17, 19, 21, 22, 23, 25. gcs-web.com-hosted IR (S-25-06) serves detail pages cleanly. For financials, prefer EDGAR over the IR site entirely |
Alphabet IR (abc.xyz) |
S-16-07 | every automated retrieval failed; Alphabet's Q2 2026 advertising revenue is absent from sector 16 as a result | EDGAR 10-Q via data.sec.gov → /Archives/ |
| Visa Onchain Analytics | S-21-12 | JS-rendered dashboard | Headless browser, or Allium data |
| Tether / Circle transparency dashboards | S-21-15 | both JS-rendered, no figures returned | Headless browser for live balances, or use on-chain supply instead |
| Verra registry | S-20-21 | registry.verra.org is a JS application |
Berkeley Voluntary Registry Offsets Database (S-20-09) bulk, or a registry API |
| Banco Central do Brasil Pix | S-07-24 | statistics page JS-rendered | BCB open-data API endpoints |
| HKMA / SFC | S-21-23 | empty JS bodies or 404 on every attempt | Headless browser or the RSS endpoint |
| AAR weekly rail traffic | S-09-15 | JS dashboard | The weekly PDF, not the dashboard |
| NY Fed labour market for recent graduates | S-22-01 | interactive charts JS-rendered | Locate the backing data files; fall back to the quarterly text |
| JNTO Japan inbound tourism | S-25-20 | jnto.go.jp 404s / client-side charts |
Japan Tourism Marketing Co. republication (registered as the working route) |
| Roblox IR | S-15-03, S-17-15 | JS-rendered; revenue, bookings and DAU not retrieved | EDGAR 8-K and shareholder-letter exhibits |
Class C — proxy 403 / 429 / rate-limited
| Source | ID | Blocker | Recommended route |
|---|---|---|---|
| CISA KEV JSON | S-04-02 | 403 to generic HTTP clients through some proxies | Browser user-agent; the JSON and CSV feeds are published for machine consumption |
| FDIC Quarterly Banking Profile | S-07-01 | QBP landing pages and PDF returned 403 | FFIEC CDR bulk download or the FDIC BankFind API. /news/press-releases/ fetched fine |
| CBP trade releases | S-12-10 | cbp.gov 403 to automated fetch |
Newsroom RSS rather than page fetches |
| Monetary Authority of Singapore | S-21-19 | 403 on every attempt | RSS / subscription email |
Federal Register /documents/search and /api/v1/ |
S-21-20 | search robots-disallowed; API PROXY_REJECTED 403 | govinfo bulk; direct document URLs (these work) |
| TSA daily checkpoint numbers | S-25-15 | tsa.gov 403 on every path, including with a browser UA |
Unblocked egress, or a mirrored dataset. Validate FAA/DOT (S-25-25) alongside |
| World Bank State and Trends of Carbon Pricing 2026 | S-20-10 | 403 through the research proxy | The Carbon Pricing Dashboard itself (instrument counts 47/40/34 were obtained) |
| UN Tourism Barometer | S-25-22 | 403; news archive only pre-2025 | National statistics offices; Eurostat tourism (S-25-10) |
| HESA (UK), OECD | S-22-19, S-22-20 | 403 / every landing pattern 404 — no OECD figure appears in sector 22 | HESA open data; OECD SDMX API rather than landing pages |
| ICE SEVIS / State Dept | S-22-09 | both 403 | SEVIS by the Numbers archived PDFs; State Dept FOIA library |
| FSIS | S-19-13 | 403 on news and proposed rules | Federal Register (S-19-12) — the reliable substitute for FDA and FSIS pages that block |
| Artificial Analysis model pricing | S-01-10 | /models endpoint 429 repeatedly |
API or licensed feed. No current independent price-per-capability series was obtained at all |
| Assemblée nationale | S-23-17 | aggressive 429 | Space requests; back off |
| USDA FAS agricultural trade | sector 19 | 403 | GATS / ESR bulk downloads |
Class D — URL-pattern drift and archive 404s
| Pattern | Examples | Rule |
|---|---|---|
| Deep-linked press releases 404 once archived | Eli Lilly (S-06-10), Blackstone (S-11-14), The Trade Desk (S-16-20), Booking (S-25-08), Hilton (S-25-24), NIKE (S-23-07), Symbotic (S-13-07) | Never store a constructed permalink. Resolve from the index or feed, store the resolved canonical URL and a content hash |
| Path-form traps | FCA /news/press-releases 404s, /news works (S-07-13); Fed FEDS Notes year indexes use .html while most Fed pages use .htm (S-21-06); BIS work.htm / othp_list.htm 404 (S-07-10, S-21-07); SEC litigation-release path changed form (S-21-03) |
Maintain a per-publisher path map with a last-verified date; re-verify on every 404 rather than failing the source |
| Dated file URLs | ESMA MiCA register CSV at a dated /sites/default/files/ path (S-21-08) |
Resolve the dated path from the landing page each run; never hard-code |
| Mixed fetchability within one agency | FDA: FSMA 204 and front-of-package pages fetch, constituent-update and food-chemical-safety paths do not (S-19-11) | Per-path health checks, not per-domain |
General retrieval playbook (from the macro brief, confirmed by 25 sectors)
- SEC:
data.sec.gov/submissions/CIK##########.json→/Archives/. Declared UA, ≤10 req/s. - Corporate IR: skip the index; fetch the release detail page or the PDF.
- Government statistics: direct file paths work (e.g. the Census e-commerce PDF at
census.gov/retail/mrts/www/data/pdf/ec_current.pdf); navigation pages often do not. - Federal Register is the universal substitute for any US agency page that blocks.
- Check the publication date of every result before using it. Date laundering in search results is systematic.
- Treat robots.txt as binding. Where a source is robots-disallowed, the route is an alternative endpoint, an official API, or a licence — never a workaround.
6.11 Monitoring design
| Tier | n | Cadence | Composition |
|---|---|---|---|
tier1_daily |
200 | ≤24h | 53 regulatory filings, 45 government data, 24 trade publications, 24 company IR, 11 earnings, 11 press releases, 7 pricing |
tier2_weekly |
246 | ≤7d | 44 government data, 36 trade publications, 33 regulatory/company IR, 22 market reports |
tier3_monthly |
157 | ≤30d | 41 government data, 33 market reports, 19 surveys, 14 academic |
Engineering allocation follows from three counts:
- 386 sources (64%) can be polled through an API or RSS today — build this first; it is
most of the value at the lowest cost. 317 records already carry a
feed_url. - 45 tier-1 sources have neither an API nor RSS. These are the highest-value manual surfaces in the registry and should get purpose-built fetchers with per-path health checks.
- 217 sources (36%) have neither. For these, a content-hash diff on a stable URL is the correct primitive: detect change, then parse — rather than parsing on every poll.
Cross-cutting requirements: deduplicate by endpoint and fan out by sector (§6.1); carry the
legal_use_restrictions field into the fetcher as an executable policy (rate limits,
declared user agents, attribution requirements, no-redistribution flags) rather than as
documentation; and record every fetch failure against the source record so the access-blocker
register in §6.10 stays current instead of becoming a snapshot of 2026-09-15.
6.12 The rule, one more time
No scraping behind logins. No bypassing access controls. No paywall circumvention. No ToS violations. Official APIs, RSS, public pages and licensed data only. Where that leaves a hole, the hole is published as a finding with the source registered so the gap is auditable — which is why S-20-25 (California Air Resources Board) sits in the registry as "Tier A source, zero coverage" rather than being quietly dropped.
- Source artifact
- 01-frameworks/16-source-registry.md
- Corpus date
- 15 September 2026
- Prepared for this site
- 16 September 2026
- Site publication
- 18 September 2026
- Verification
- Inherited; not fully rechecked