SASIGNAL ATLASCross-industry intelligence / Research desk
SIGNAL ATLAS / RESEARCH DESK

6 — Data Sources and Source-Quality Framework

Method · Cross-industry · Original Phase 1 research

6 — Data Sources and Source-Quality Framework

Phase 1 deliverable · research date 2026-09-15 · registry snapshot: 603 records


Part 1 — The source map

6.1 What the registry is

/home/claude/research/data/consolidated/sources.json holds 603 source records — the sources a Phase 2 pipeline would actually poll, not a bibliography. Each was registered by the sector analyst who used it, against the schema in §7 of the research contract, and each carries 21 fields including reliability, bias, update frequency, access method, cost, geographic coverage, historical depth, citation value, legal use restrictions and monitoring priority.

Measure Value
Records 603
Sectors covered 25 (22–25 records each; min 22 at sectors 02, 10, 13, 18, 19; max 25)
Distinct publishers 494
Distinct domains 454
Distinct (name, publisher) pairs 597
Distinct source types 24 of the schema's 27 enumerated values, plus one off-schema value (trade_org, 1 record, S-05-17)
Machine-readable (API or RSS) 386 (64.0%)
Neither API nor RSS 217 (36.0%)
Non-null feed_url 317 (52.6%)
Free at point of use 497 (82.4%)
Requiring payment of any kind 106 (17.6%)

Deliberate duplication is a feature of the registry and a requirement on the pipeline. 603 records resolve to 454 domains because the same primary source was independently registered by several sector analysts: www.sec.gov, data.sec.gov and efts.sec.gov together account for 33 records across 22 of the 25 sectors; www.census.gov 14; www.bls.gov 11; www.federalreserve.gov 8; www.federalregister.gov 7. Phase 2 must deduplicate by endpoint for polling and fan out by sector for routing — one EDGAR crawler, 22 sector subscriptions — or it will hit the SEC's 10-requests-per-second fair-access limit 33 times over with 33 separate user agents.

6.2 Composition

By source type

Type n % Typical reliability API RSS Default monitoring tier
government_data 130 21.6% high (121/130) 43% 44% mixed (45 t1 / 44 t2 / 41 t3)
regulatory_filing 92 15.3% high (90/92) 52% 76% tier1 (53 t1)
trade_publication 73 12.1% medium (43/73) 1% 82% tier2 (36 t2 / 24 t1)
company_ir 63 10.4% high (59/63) 3% 48% tier1–2
market_report 58 9.6% medium (34/58) 14% 29% tier3 (33 t3)
earnings 30 5.0% high (30/30) 20% 67% tier2
survey 27 4.5% mixed (12 high / 15 med) 11% 41% tier3 (19 t3)
academic 25 4.1% high (24/25) 20% 76% tier2–3
press_release 24 4.0% mixed (12/12) 0% 50% tier1–2
public_dataset 19 3.2% high (12/19) 37% 11% tier2
standards_body 14 2.3% high (13/14) 7% 50% tier3
startup_db 11 1.8% medium (10/11) 64% 55% tier2
pricing 10 1.7% high (8/10) 30% 10% tier1 (7/10)
trade_data 6 1.0% high (4/6) 17% 0% tier1–2
newsletter 4 0.7% mixed 0% 100% tier2
developer_activity 3 0.5% high 100% 33% tier1–2
procurement 3 0.5% high (3/3) 33% 33% tier1
app_store 3 0.5% medium 67% 0% tier2
web_traffic 2 0.3% mixed 50% 50% tier1–2
vc_announcement 2 0.3% medium 50% 100% tier1–2
job_postings 1 0.2% medium 0% 0% tier2
open_source 1 0.2% high 0% 0% tier2
trade_org (off-schema) 1 0.2% medium 0% 0% tier3
conference 1 0.2% low 0% 0% tier3
patent 0 absent
search_trends 0 absent
social 0 absent
community 0 absent

The top four types are 358 records — 59.4% of the registry. The registry is overwhelmingly a primary-document registry: government statistics, regulatory filings, company disclosure and named-byline trade press. That is the correct shape for a platform whose evidence cap punishes weak sourcing, and it is a direct consequence of the tiering rules in §6.5. It is also a coverage bias that Phase 2 must correct deliberately (§6.4).

By reliability, cost, access and cadence

Reliability n % Cost n %
high 450 74.6% free 497 82.4%
medium 152 25.2% freemium 78 12.9%
low 1 0.2% paid_high 17 2.8%
enterprise 9 1.5%
paid_low 2 0.3%
Access method n % Update frequency n % Monitoring tier n %
public_page 336 55.7% quarterly 130 21.6% tier2_weekly 246 40.8%
rss 101 16.7% daily 124 20.6% tier1_daily 200 33.2%
api 91 15.1% irregular 111 18.4% tier3_monthly 157 26.0%
bulk_download 62 10.3% monthly 100 16.6%
licensed 12 2.0% weekly 69 11.4%
email 1 0.2% annual 48 8.0%
realtime 20 3.3%
semi-annual 1 0.2%

Cross-cuts that matter for engineering:

  • 404 of 603 records (67%) are both high reliability and free. The intelligence product is overwhelmingly buildable on free primary sources. The licensing budget in §6.9 exists to close specific, named holes — not to buy the backbone.
  • public_page at 55.7% is the pipeline's central problem. More than half the registry must be fetched as HTML, which is exactly the population that the access-blocker register (§6.10) shows is most fragile.
  • 45 of the 200 tier-1 daily sources have neither an API nor an RSS feed. These are the highest-frequency, highest-value, least-automatable sources in the registry and they should be the first engineering allocation.
  • Only 12 records are licensed, and 10 of those are market_report. Licensing is concentrated, which is why §6.9 can be a short, prioritised list rather than a subscription sprawl.

By geography

Coverage bucket n %
Global 252 41.8%
US-only 249 41.3%
EU / Europe 40 6.6%
Asia 29 4.8%
Other / regional 33 5.5%

41.3% US-only is the registry's largest structural bias, and the "global" bucket overstates coverage: 34 of the 252 global records carry an explicit skew in the field itself ("Global, US-weighted", "Global, US-skewed", "Global with strong Taiwan, Korea and China coverage"). The contract's non-English requirement did produce real regional depth — 48 records from non-US publishers including TSMC, SK hynix, Samsung, TrendForce, DigiTimes, Nikkei Asia, the Taiwan Stock Exchange MOPS (S-03-23), Korea Customs/MOTIE (S-03-24), CAAM (S-10-02), CPCA via CnEVPost (S-10-03), Gasgoo (S-10-14), China GACC (S-09-19), NPCI (S-07-23), VDMA (S-09-17), Eurostat (S-09-18), ENTSO-E (S-05-11), KOCCA (S-14-19), NPPA (S-15-06), JFTC (S-15-22), JARA (S-13-16), CRIA/MIIT (S-13-17), CONAB (S-19-20), the Federation of the Swiss Watch Industry (S-23-22) — but that is 8% of the registry against sectors whose centres of gravity are frequently outside the US.


6.3 Category-by-category evaluation

Each category below is assessed on the twelve attributes the brief requires. The compact table carries the quantitative attributes; the prose carries bias, data quality, historical depth, citation value and legal restrictions, which do not compress.

Summary table

Category (registry type) n Reliability Update Access Cost API/RSS Geo Citation value
Government publications (government_data) 130 high monthly 50, daily 18, weekly 17 public_page 54, bulk 33, api 30 free 126 43% / 44% US-heavy, strong EU/Asia tail high
Regulatory filings (regulatory_filing) 92 high daily 29, irregular 32, realtime 13 api 43, public_page 35 free 91 52% / 76% US-dominant high
Trade publications (trade_publication) 73 medium daily 50 rss 41, public_page 30 free 45, freemium 21, paid 6 1% / 82% global medium
Company sites & IR (company_ir) 63 high quarterly 45 public_page 51 free 63 3% / 48% global high
Market reports (market_report) 58 medium quarterly 17, annual 13, monthly 12 public_page 40, licensed 10 free 29, paid/ent 15 14% / 29% global medium
Earnings (earnings) 30 high quarterly 30 public_page 26 free 30 20% / 67% global high
Consumer/market surveys (survey) 27 mixed annual 12, monthly 6 public_page 22 free 19 11% / 41% US-heavy medium
Academic papers (academic) 25 high weekly 6, irregular 7 public_page 11, rss 7, bulk 5 free 20 20% / 76% global high
Press releases / product announcements (press_release) 24 mixed irregular 14 public_page 18, rss 6 free 24 0% / 50% global medium
Public datasets (public_dataset) 19 high daily 8, quarterly 4 bulk 8, public_page 9 free 14 37% / 11% global high
Standards bodies (standards_body) 14 high irregular 9 public_page 11 free 13 7% / 50% global/EU high
Startup databases (startup_db) 11 medium 10/11 daily 5, quarterly 4 api 4, public_page 5 freemium 8, paid 2 64% / 55% global, US-weighted medium
Pricing (pricing) 10 high daily/weekly/monthly public_page 8 free 6, freemium 4 30% / 10% global high
Import/export (trade_data) 6 high monthly 6 public_page 4 free 4 17% / 0% US, CN, CH, EU high
Newsletters (newsletter) 4 mixed weekly 3 rss 3 freemium 3 0% / 100% global medium
Developer activity (developer_activity) 3 high daily 2 api 2 free 3 100% / 33% global medium
Procurement (procurement) 3 high daily 2 public_page 2, api 1 free 3 33% / 33% US federal high
App stores (app_store) 3 medium realtime/weekly/monthly public_page 3 free 1, freemium 2 67% / 0% global medium
Web traffic (web_traffic) 2 mixed realtime/monthly api 1, public_page 1 freemium 1, free 1 50% / 50% global / US medium
VC announcements (vc_announcement) 2 medium weekly 2 rss 2 freemium 2 50% / 100% global medium
Job postings / hiring (job_postings) 1 medium weekly public_page free 0% / 0% US medium
Open source (open_source) 1 high irregular public_page free 0% / 0% global medium
Conferences (conference) 1 low annual public_page paid_high 0% / 0% global low
Patents 0
Search trends 0
Social 0
Community 0
M&A 0 as a type covered via regulatory_filing, earnings, trade_publication

Government publications — 130 records, the registry's backbone

Named: EIA Short-Term Energy Outlook and Electric Power Monthly (S-05-01, S-05-02 — API + RSS, monthly, back to 1990/1997); US Census Value of Construction Put in Place C30 (S-09-01, S-18-01, bulk, back to 1993); Census Advance Monthly Retail Trade (S-12-02, S-19-09); Census Business Trends and Outlook Survey with AI questions (S-01-08, S-09-21, weekly, back to 2022); BLS CPI (S-19-01, back to 1913), PPI (S-18-05, back to 1947), Employment Situation (S-18-04, S-22-02, back to 1939/1948), JOLTS (S-22-04, back to 2000); Federal Reserve G.17 Industrial Production (S-09-03, back to 1919) and H.8 (S-07-02, back to 1973); FRED's Indeed software-development job postings index (S-02-04) and BVP Emerging Cloud Index (S-02-05); CISA advisories and the KEV catalog (S-04-01, S-04-02); FDA Novel Drug Approvals (S-06-01) and ClinicalTrials.gov (S-06-05, API, back to 1999); USAspending.gov (S-08-03, API, back to FY2008); ERCOT large-load interconnection queue (S-05-06, weekly, API + RSS); CAAM (S-10-02) and CPCA via CnEVPost (S-10-03); FAO Food Price Index (S-19-07, back to 1990); USDA WASDE (S-19-05, monthly since 1973); EPA Greenhouse Gas Reporting Program (S-20-11, bulk, back to 2010); IPEDS (S-22-05).

Reliability high (121 of 130). Bias: not commercial, but definitional and political — the statistical agency decides the category, and the category can hide the story. Phase 1 confirmed two traps that any pipeline must encode: US Census reports data-centre construction inside the "Office" category, which is why Census office construction was +21.3% y/y in the middle of an office-distress narrative; and core CPI 2.4% (Aug 2026) against core PCE ~3.3%, an unusually wide gap in the unusual direction. Never treat "inflation" as one number; always name the index. Data quality high; revisions are published and must be ingested as revisions rather than as new facts. Historical depth is the category's decisive advantage — several series run 50–100 years, which is what makes calibration_basis: base_rate possible at all (§9.7). Citation value high. Legal: US federal works are public domain; the binding constraints are operational, not legal — SEC requires a declared User-Agent with contact details and 10 req/s; most agencies expect attribution. API/RSS: 43% / 44%, the best of any large category, and 30 records already expose a real API.

Regulatory filings — 92 records, the highest-tier-1 concentration

Named: SEC EDGAR, registered 33 times across 22 sectors (S-01-01, S-07-04, S-13-03, S-21-01 and others) — full text back to 2001, filings back to 1993/94, plus the XBRL company-facts API (S-19-16, S-13-03); Federal Register with API (S-09-13, S-19-12, S-21-20, back to 1994 in full text); FERC eLibrary (S-05-04, daily, API + RSS, back to 1981); NRC ADAMS (S-05-13); NHTSA Standing General Order incident data (S-10-09); California DMV autonomous-vehicle disengagement reports (S-10-07) and CPUC quarterly AV reports (S-10-06); FCC IBFS satellite filings (S-08-14); ESMA MiCA interim register (S-21-08, weekly CSVs); BIS export-control actions (S-03-12); EUR-Lex Official Journal L series (S-20-01, back to 1952); European Parliament Legislative Observatory (S-20-02); Taiwan Stock Exchange MOPS (S-03-23 — mandatory monthly revenue filings, the best high-frequency semiconductor disclosure anywhere); CourtListener/RECAP (S-16-22, S-17-14 — free API, realtime).

Reliability high (90 of 92) — this is the Tier A core. Bias: none in the record; substantial in what is absent. Confidential draft S-1 submissions do not appear until public filing, so absence of a filing is not absence of an IPO process (S-01-01's own bias note). Sealed filings are absent from RECAP by definition. Historical depth excellent. Citation value the highest in the registry. Legal: public domain or open-licence; rate limits and declared user agents are mandatory. API/RSS 52% / 76% — the most machine-readable category. Access caveat: EDGAR's browse-edgar CGI is robots-disallowed and efts.sec.gov 403s through some proxies; the working route is data.sec.gov/submissions/CIK##########.json → fetch under /Archives/ (§6.10).

Company sites, IR, product announcements and press releases — 87 records combined

Named IR: NVIDIA (S-01-02, S-03-07), the hyperscaler set (S-01-03), TSMC with monthly revenue disclosure (S-03-01), SK hynix (S-03-03), Samsung (S-03-04), Micron (S-03-05), Broadcom (S-03-08), ASML (S-03-09), SpaceX (S-08-05 — first IR disclosure from Q2 2026), PJM Inside Lines (S-05-05), Netflix (S-14-03), Warner Bros. Discovery (S-14-07), Roblox (S-15-03, S-17-15), Meta (S-17-04), Snap (S-17-09), Reddit (S-17-20), Pinterest (S-17-16), Richemont (S-23-03), L'Oréal (S-23-04), Estée Lauder (S-23-05), Inditex (S-23-24), Kering (S-23-25), Spotify 6-K via EDGAR (S-24-08), Deezer (S-24-10). Named press-release/product-announcement sources: Anthropic newsroom (S-01-04), OpenAI index (S-01-05), TSMC Latest News (S-03-02), FDA Press Announcements (S-06-02), Apple Developer News and regional fee schedules (S-15-01, S-17-03), YouTube Official/Creator blog (S-17-10), TikTok Newsroom (S-17-23), USDA (S-19-10), Suno blog (S-24-17), IATA Press Room (S-25-01), aggregate media-sector issuer RSS (S-14-24).

Reliability high (59 of 63 IR records) for what is disclosed; bias is the issuer's own framing, and vendor product announcements are marketing claims by definition and must carry the label. Data quality high for audited figures, low for anything in a slide footnote. Historical depth 10–30 years for most listed issuers. Legal: copyright to the issuer; press material is normally citable with attribution; do not redistribute images or full text. API/RSS 3% / 48% — and this is the category's operational weakness: IR index pages are overwhelmingly JavaScript-rendered and fail to a plain fetch, while individual press-release detail pages and PDF links fetch reliably. This was confirmed independently in sectors 13, 16, 17, 19, 21, 22, 23 and 25. Route accordingly (§6.10).

Earnings — 30 records, quarterly, uniformly high reliability

All 30 are reliability: high and all 30 are quarterly — the only category in the registry with perfect internal consistency. Named: Micron (S-03-05), NVIDIA (S-03-07), Broadcom (S-03-08), ASML (S-03-09), GE Vernova orders (S-05-12), Tencent and NetEase interim reports (S-15-13), The Trade Desk (S-16-20), Meta (S-17-04), Snap (S-17-09), Pinterest (S-17-16), Reddit (S-17-20), NIKE (S-23-07), adidas (S-23-08), lululemon (S-23-10), Spotify (S-24-08). Bias: non-GAAP presentation and segment definition are chosen by the issuer; segment boundaries move. Citation value high; legal as for IR. 20% API / 67% RSS — the RSS route is the right default, with EDGAR as the authoritative backstop.

Academic papers — 25 records

Named: arXiv quant-ph plus Nature/Science quantum error-correction coverage (S-03-21, API + RSS); Nature Medicine (S-06-07); Epoch AI data hub (S-01-09, API + bulk, models database back to 1950); LBNL Queued Up interconnection-queue dataset (S-05-08); Penn Wharton Budget Model effective tariff rates (S-09-12); Rhodium Group China research (S-10-19); Duke Nicholas Institute (S-05-24); Cloud Security Alliance (S-01-23); Future of Privacy Forum legislative trackers (S-17-18). Reliability high (24/25); bias — preprints are unreviewed and arXiv volume tracks incentives as much as progress; think-tank "academic" output carries funder positions and must be read as such. Update frequency irregular to weekly. Historical depth deep. Citation value high. Legal: licences vary — arXiv per-paper licences, many journals are closed-access; do not circumvent journal paywalls; use the preprint or the abstract. 20% API / 76% RSS.

Patents and patent databases — 0 records. The registry's single largest gap.

The schema enumerates patent as a source type. No sector registered one. This is a material omission rather than a judgement that patents do not matter: §7 of the Phase 1 methodology weights leading indicators — hiring, patents, procurement, supply-chain shifts — above lagging ones, and the patent leg is simply absent from the source layer. Patents were referenced in dossier prose (ASML and export-licensing context in sector 03; ISO/TC 299 work items in 13) but never as a monitorable feed.

Phase 2 must register, at minimum: USPTO PatentsView and the USPTO Open Data Portal (bulk + API, free, US, back to 1976, public domain); EPO Open Patent Services / Espacenet (API, freemium, global families, attribution and rate limits); WIPO PATENTSCOPE (public pages, global PCT); Google Patents BigQuery public dataset (bulk, free tier). Expected attributes: reliability high, update weekly, access api/bulk, cost free–freemium, geographic coverage global with well-known filing-jurisdiction bias, data quality high but noisy as a trend signal (filing volume tracks legal strategy and subsidy regimes as much as invention), historical depth decades, citation value high, legal restrictions minimal for US, attribution and rate limits for EPO.

Search trends — 0 records

Also enumerated in the schema and also unregistered. Partly defensible: the contract is explicit that search interest is an input to velocity only and that velocity is 7.5% of the composite, so a search-trends feed can never be load-bearing. But the absence means the platform currently has no instrument for the fad diagnostic — attention up, adoption flat — which is how §1.3 defines a fad. Phase 2 should register Google Trends (public page / unofficial API, free, global, rolling, relative not absolute values, low data quality as a level, useful only as a within-series delta) and Wikipedia Pageviews (official API, free, bulk, back to 2015, genuinely absolute counts, high data quality, public-domain licence) — and mark both citation_value: low so they cannot be laundered into evidence.

Social and community — 0 records as types; 12 adjacent

No social or community source is registered. The adjacent coverage is: the EU DSA Transparency Database (S-17-24 — European Commission, daily bulk download, one common schema across TikTok, Meta, Snap, Pinterest and X, tier1_daily, free, high reliability); platform newsrooms as press releases (S-17-10 YouTube, S-17-23 TikTok); platform IR (Meta S-17-04, Snap S-17-09, Reddit S-17-20, Pinterest S-17-16); Ofcom Online Safety enforcement (S-17-01); the Australian eSafety Commissioner (S-17-05); Pew Research (S-17-08, S-22-15); and one genuine community artefact registered as a public dataset — videogamelayoffs.com (S-15-08, community-maintained, medium reliability, irregular).

This is the right shape and the wrong volume. Platform-reported regulatory data is Tier A–adjacent; raw social volume is not evidence of anything under this contract. The DSA database was flagged by sector 17 as "the best structured source in its sector" and is still unexercised. Phase 2 should wire it first among social sources. Community forums (subreddits, Discord, HN) should be registered only with citation_value: low and reliability: low, used for discovery and never for support — and only via official APIs under their terms, never by scraping authenticated surfaces.

Startup databases, VC announcements and M&A — 13 records + no dedicated M&A type

Named: Crunchbase, registered seven times across sectors 01, 02, 04, 07, 08, 11 and 22 (S-01-24, S-02-14, S-04-24, S-07-15, S-08-20, S-11-07, S-22-13 — API + RSS, freemium to paid_high, global, US-weighted); Carta Data Desk (S-11-03 — free, quarterly, high reliability, US-primarily, the only free public source for down-round rates, bridge rounds, stage-level round counts, AI vs non-AI valuation splits and fund-level DPI by vintage); Dealroom (S-11-17, Amsterdam, API); Tracxn (S-11-18, Bengaluru); CRETI proptech funding (S-18-19); CDR.fyi (S-20-06); Sightline Climate / CTVC (S-20-16, RSS, weekly).

Reliability medium for 10 of 11 startup_db records — the only category where medium is the norm, and correctly so. Bias: coverage is a function of who self-reports; Crunchbase and KPMG differ by $50.4bn on the same half-year of global VC ($510B vs $560.4B) and by 5,000+ vs 8,440 on Q2 deal count, a definitional difference in deal inclusion rather than an error. Carta's sample is self-selected: US-weighted, skewed to smaller and earlier-stage companies and sub-$100M funds. Data quality medium; historical depth ~2018 onward for Carta, mid-2000s for the commercial databases. Citation value medium — always name the tracker. Legal: redistribution of underlying datasets prohibited; free news tiers citable with attribution. API 64% / RSS 55% — the best-instrumented small category.

M&A has no dedicated source type. It is covered through regulatory_filing (8-K Item 2.01, S-4, HSR-driven disclosures — Agility Robotics' S-4, filed 2026-09-04, CIKs 0002074973 and 0001727116, is the worked example), earnings, DOJ Antitrust press releases (S-14-15), state AG filings (S-14-16), and trade press. This is adequate for verification and inadequate for discovery: Phase 2 should add an M&A-typed feed keyed on EDGAR form types and antitrust dockets rather than buying a deal database.

Job postings and hiring — 1 record typed, ~6 functional

Named: the only job_postings record is S-02-19 (TechCrunch AI layoff tracker plus US state WARN notices, medium reliability, weekly, US). Functionally the hiring signal is carried by government data: FRED's Indeed software-development job postings index (S-02-04, API, daily, baseline Feb 2020), BLS JOLTS (S-22-04, back to 2000), BLS Employment Situation construction detail (S-18-04), BLS Occupational Outlook and Employment Projections (S-02-06), and NACE first-destination surveys (S-22-25).

Given that hiring is one of the four leading indicators the methodology explicitly weights above media coverage, one typed record is under-provisioned. The structural finding that came out of hiring data in Phase 1 was among the strongest in the programme — US construction unemployment at a record-low 3.1% despite a soft market, with fewer than half of US metros adding construction jobs year over year, confirming data-centre crowd-out from the labour side. Phase 2 should register state WARN feeds directly (free, bulk, US, high reliability, legally clean) and, if a commercial postings panel is required, treat it as a licensing decision with an explicit bias note: job-posting panels measure advertised demand, which diverges from hiring in both directions.

App stores and product launches — 3 records

Named: Steam AI-content disclosure fields and store metadata (S-15-09 — Valve, free, realtime, tier1_daily, high reliability, global PC); Appfigures Insights (S-17-11) and Appfigures/Sensor Tower public releases (S-15-25), both freemium with partial API. Bias: store-side metadata is issuer-declared (Steam's AI disclosure is self-reported); third-party estimators model downloads and revenue from sampled panels and do not publish methodology. Data quality medium for estimates, high for declared metadata. Historical depth shallow. Citation value medium. Legal: store ToS prohibit bulk scraping — use official APIs and published reports only. Critical context: app-store economics fragmented by jurisdiction inside nine months — the flat 30% commission was dismantled or altered across the US, EU, Japan, China, Korea, UK and Brazil; Apple's EU 5% CTC applies from 2026-10-01 and China 25%/12% from 2026-03-15. Any app-economy model must now be jurisdiction-specific, which makes the fee-schedule pages themselves (S-15-01, S-15-02) a higher-value monitoring target than download estimates.

Developer activity and open source — 4 records

Named: Hugging Face Hub API and blog (S-01-18 — API + RSS, daily, tier1_daily, free, high reliability); package-registry and developer telemetry across npm, PyPI and GitHub (S-02-22 — API, daily, medium reliability, token-gated and rate-limited); the agentic-commerce protocol documentation from Stripe, OpenAI and Google (S-12-20); OWASP GenAI Security Project (S-04-15, the sole open_source record). The Model Context Protocol specification (S-02-07) sits under standards_body and functions as the integration tripwire for enterprise software.

100% of developer_activity records expose an API — the highest rate in the registry. Bias: registry download counts are heavily inflated by CI systems and mirrors and are close to useless as adoption measures without de-botting; GitHub stars are a promotion metric. Data quality high for counts, low for interpretation. Legal: GitHub API requires a token and rate-limits; npm and PyPI publish open datasets. Citation value medium — a real leading indicator, easily gamed, never sole support.

Newsletters — 4 records, all RSS

Named: SemiAnalysis (S-03-18 — freemium, weekly, strong China and Taiwan, subscriber-only for the substantive work); GameDiscoverCo (S-15-10 — high reliability, weekly, tier1_daily); GameDev Reports (S-15-18); aggregated law-firm regulatory advisories from Arnold & Porter, Latham and peers (S-06-25 — high reliability, free, irregular). 100% RSS availability, which makes them cheap to ingest. Bias: single-author newsletters carry a single analytical prior and frequently a commercial one; law-firm advisories are marketing for a practice group but are unusually reliable on what a rule says and when it takes effect — which is exactly the thing vendor marketing gets wrong. Citation value medium; use as pointers to primaries, not as primaries.

Trade publications — 73 records, the discovery layer

Named, high value: Utility Dive (S-05-15), BioPharma Dive (S-06-13), Retail Dive (S-12-22), Construction Dive (S-18-20) — the Industry Dive stable, free, daily, RSS, named bylines; The Register (S-02-12); TechCrunch (S-02-13); The Robot Report (S-13-05); SpaceNews (S-08-18); Breaking Defense (S-08-19); AdExchanger (S-16-11); Digiday (S-16-12); Data Center Frontier (S-18-21); DatacenterDynamics (S-01-12); PocketGamer.biz (S-15-16); Game Developer (S-15-17, back to 1994 as Gamasutra); CnEVPost (S-10-21) and Gasgoo (S-10-14) for China; South China Morning Post Tech (S-01-17); Nikkei Asia (S-03-17, metered); DigiTimes (S-03-15, subscription — and the likely origin of much uncredited Taiwan supply-chain reporting); TrendForce (S-03-14); STAT News (S-06-14) and Endpoints (S-06-15), both paywalled; Reuters Health and Pharmaceuticals (S-06-23, enterprise); Courthouse News antitrust (S-16-21).

Reliability medium for 43 of 73 — correctly graded down. 82% RSS availability makes this the cheapest high-volume ingest in the registry, and the reason trade press is the discovery layer rather than the evidence layer. Bias: advertiser and event revenue from the industry covered; Industry Dive, Informa and Access Intelligence all run conferences for the sectors they report on. Legal: copyright; headlines and short quotation with attribution only; no full-text redistribution and no paywall circumvention. Syndication is the central risk and is handled in §6.7.

Conferences — 1 record, and it is the registry's only low reliability source

S-16-24 (Cannes Lions / Advertising Week / Possible programmes and disclosures) is the sole conference record, reliability: low, citation_value: low, cost: paid_high, and its bias note is the correct general rule for the whole category: "Pay-to-present environments. Almost every quantitative claim made on stage is vendor-sponsored and unaudited. Useful as a leading indicator of narrative, explicitly not of adoption." Conference programmes are a legitimate narrative instrument — they show what a sector has decided to talk about six months ahead — and must never support a factual claim.

Consumer and industry surveys — 27 records

Named: Census Business Trends and Outlook Survey (S-01-08, S-09-21, S-13-14 — weekly, API, the only official US AI-adoption series); ISM Manufacturing PMI (S-09-07, back to 1948); Verizon DBIR (S-04-06); Stack Overflow Developer Survey (S-02-08); Pew Research (S-17-08, S-22-15); NAM Manufacturers' Outlook (S-09-16); Modern Materials Handling / Peerless intralogistics robotics surveys (S-09-24, S-13-19); WFA (S-16-15); CDP disclosure and scores (S-20-18); NACE (S-22-25); GBTA Business Travel Index (S-25-17, paywalled data cube).

Reliability is genuinely mixed (12 high / 15 medium) and this is the category where the fact vs estimate line matters most. Under §1.3, survey-reported intent is an estimate; transaction, traffic and usage data are fact — and the two diverge constantly, with travel the worst offender. Bias: trade-association surveys have a membership interest in the narrative; panel composition is rarely disclosed; year-over-year change within the same panel is the usable signal, levels are not. The canonical trap from sector 25: GBTA's 2026 release shows spend +7.2% against trips +1.3%; quoting only the first is the sector's most common analytical error. Legal: press releases quotable; data cubes licensed.

Market reports — 58 records, the most problematic category

Named, credible: SEMI WWSEMS (S-03-10), WSTS (S-03-11, back to 1986), IFR World Robotics (S-13-01 free releases / S-13-02 licensed reports), Nielsen The Gauge (S-14-01), Circana (S-15-11), BloombergNEF (S-05-19, S-10-13), WARC (S-16-17), EMARKETER (S-12-15), Comscore (S-14-12), Trepp (S-07-19, S-18-11), Green Street CPPI (S-18-12), PitchBook (S-07-21), Coresight (S-12-17), SNE Research (S-10-12), Madison & Wall (S-16-18), Grid Strategies (S-05-23), Tax Foundation tariff tracker (S-09-25), Marsh Global Insurance Market Index (S-04-09), KPMG Venture Pulse (S-07-16), Chainalysis (S-04-08).

Reliability medium for 34 of 58, tier3_monthly for 33 of 58, and 15 of the 28 strictly-paid records in the whole registry (paid_low + paid_high + enterprise) sit here. The category is bimodal: a small set of methodologically serious industry-body series (SEMI, WSTS, IFR, WSTS-derived), and a large set of commercial forecasters whose free tier is marketing for the paid product.

The structural pattern to encode: headline free, detail paid. SEMI publishes the $40.53bn global equipment total and withholds the regional table where China's trajectory is visible. IFR publishes installations, stock and density and withholds end-use, application, vendor-level and service-robot data. Coresight publishes a weekly summary and withholds the line items. Green Street publishes monthly percentage changes and withholds index levels. This is not an accident of pricing; it is the product design, and it means the free tier is systematically sufficient for direction and systematically insufficient for level. Trend detection can run on the free tier. Any claim about a level needs the licence or must be labelled estimate with the modeller named.

Bias: commercial interest in the market looking large; methodology usually undisclosed; forecasts revised silently. Legal: redistribution prohibited almost universally; attribution required for headlines. Citation value medium, and Tier C on sight for anything without a named methodology (§6.6).

Public datasets — 19 records, the highest-leverage underused category

Named: EU DSA Transparency Database (S-17-24 — daily bulk, free, high reliability, common schema across all designated platforms); IMF PortWatch (S-09-23 — daily bulk, API, free, ~1,400 ports and major chokepoints, AIS-derived); CVE Program cvelistV5 (S-04-22 — use the git repository, not the website; back to 1999); Berkeley Voluntary Registry Offsets Database (S-20-09 — quarterly bulk, free, global, the working substitute for the JS-rendered Verra registry); Global Energy Monitor coal plant tracker (S-05-18); SIPRI arms industry and military expenditure databases (S-08-21, annual bulk); DefiLlama (S-21-10, realtime API), RWA.xyz (S-21-09), Artemis (S-21-11); Visa Onchain Analytics (S-21-12 — free, daily, high reliability, the only published methodology that separates genuine stablecoin payments from bot, bridge, MEV and exchange-internal traffic); IPEDS (S-22-05); National Student Clearinghouse enrolment (S-22-06); ITRC data-breach reports (S-04-07); NOAA Billion-Dollar Disasters (S-20-23); NCREIF (S-18-13, licensed); Box Office Mojo (S-14-10) and The Numbers (S-14-11).

37% expose an API; 8 of 19 are bulk downloads. Bias varies sharply: NOAA's series is authoritative and has stopped; crypto dashboards are self-published by interested parties and, per sector 21, "the sector's headline metrics are manufacturable and demonstrably manufactured" — an intelligence pipeline that ingests aggregator dashboards uncritically will import fabricated data at scale. Visa's dashboard is the counter-example precisely because its adjustment criteria are disclosed (exclude addresses exceeding 1,000 transactions or $10m in 30 days, CEX flows, mint/burn and infrastructure operations). Legal: mostly open licences with attribution. This category contains the registry's best cost-to-value ratio and its least-exercised assets.

Standards bodies — 14 records

Named: NIST CSRC post-quantum migration (S-04-17), NERC Reliability Assessments (S-05-03), ENTSO-E Transparency Platform (S-05-11 — realtime API, 39 TSOs across 36 countries), Model Context Protocol (S-02-07), FSB (S-07-09), ECB digital euro rulebook (S-07-12), ILPA (S-11-19), ISO/TC 299 Robotics (S-13-20, paid_low — ISO sells its standards), Media Rating Council (S-16-10), EFRAG ESRS (S-20-03), ICVCM (S-20-07), Verra (S-20-21), 1EdTech (S-22-24), European Commission AI Act GPAI Code of Practice (S-24-24). Reliability high (13/14); update irregular (9/14). Value: standards bodies are the earliest formal signal of a regulatory trend, and they publish work items before rules exist. Legal: ISO and several others sell the standard text — budget paid_low per standard, and never republish clause text.

Procurement — 3 records, all US federal, all high reliability

DoD daily contract announcements over $7.5m (S-08-01 — RSS, daily, tier1_daily), SAM.gov contract opportunities and awards (S-08-02 — API, daily), Space Force / SDA acquisition announcements (S-08-15). Plus USAspending.gov under government data (S-08-03, API, back to FY2008). Bias: US-only, and classified programmes are absent by construction — absence of an award is not absence of a programme. Citation value high; legal: public domain. Procurement is a named leading indicator in the methodology and is currently a single-country instrument; EU TED and UK Contracts Finder are the obvious Phase 2 additions.

Import/export — 6 typed records plus substantial government-data coverage

Named: US FT-900 (S-09-04 — Census/BEA, API, monthly, full country and HS-code detail); BLS import/export price indexes (S-09-05); China GACC monthly trade statistics (S-09-19); China customs rare-earth and permanent-magnet export data mirrored by Silverado Policy Accelerator (S-13-18 — monthly, tier1_daily, high reliability, by destination country); Korea Customs / MOTIE semiconductor exports (S-03-24); OTEXA US textile and apparel trade (S-23-23 — bulk, free, tier1_daily); Federation of the Swiss Watch Industry exports (S-23-22); USITC HTS and CBP CSMS tariff guidance (S-03-13); Port of LA / Long Beach container statistics (S-09-11).

This category punched far above its weight in Phase 1. Korean monthly export values are described in the sector 03 dossier as "historically the earliest reliable turn signal in this sector." Chinese rare-earth licensing data produced the finding that Japan received zero covered rare-earth exports in July 2026, which is a robotics and industrial constraint, not only a defence one. Bias: national statistics reflect national framing; transshipment makes origin attribution unreliable, and the split between genuine relocation and transshipment in Vietnam and Mexico flows was explicitly recorded as unresolvable from available data. Legal: government data, free. 0% RSS, 17% API — this category needs bulk-download plumbing, not feeds.

Pricing — 10 records, 7 of them tier1_daily

Named: Drewry World Container Index (S-09-08, weekly), Cass Transportation Index (S-09-09, monthly), DAT spot rates and load-to-truck (S-09-10, weekly, partial API), World Bank Pink Sheet commodity prices (S-19-08, monthly bulk + API, global), IATA Jet Fuel Price Monitor (S-25-03, weekly), Kelley Blue Book average transaction price (S-10-04, monthly), OpenRouter LLM rankings (S-01-11, API, daily), software vendor pricing and packaging pages (S-02-16, tier1_daily — pricing-page diffing is the instrument for the per-seat-to-consumption shift, T-02-01), Apple and Google fee schedules (S-15-01, S-15-02), BVP cloud index multiples (S-02-21).

Highest tier-1 density of any category (7 of 10) because prices move daily and are the least-laggy observable in the registry. Bias: index construction is proprietary for Drewry, Cass and DAT; vendor list prices diverge from realised prices. Data quality high; historical depth decades for the commodity and freight series. Legal: Drewry restricts automated access via robots.txt and its full dataset is a licensing cost; World Bank is open. Gap: there is no audited CPM index for advertising in any channel, and no current independent price-per-token index — Epoch AI's inference-price dataset was last updated 2025-03-12 and Artificial Analysis's pricing table returned HTTP 429 repeatedly.

Web traffic — 2 records

Cloudflare Radar (S-04-18 — API, realtime, freemium, global with country breakdowns, high reliability) and Adobe Digital Insights retail and AI-traffic reports (S-12-16 — monthly, free, US, medium reliability, published by an interested vendor). This is thin for a category that matters, particularly as AI-referred traffic becomes a live commercial question in sectors 12, 16, 17 and 24. Bias: every commercial traffic estimator models from a panel and none publishes the panel. Phase 2 should add Cloudflare Radar's full API surface and the Wikimedia Pageviews API (genuinely counted, not modelled) and should mark all panel-derived traffic estimates estimate, never fact.


6.4 The registry's own biases — read this before trusting the map

  1. US-centricity. 41.3% US-only; the "global" bucket contains 34 records with an explicit skew in the field itself. Sectors whose centre of gravity is outside the US (semis, EV/battery, luxury, K-content, shipping) met the regional-source requirement, but sector 16 recorded "China, India, Brazil and Southeast Asia: entirely absent" and sector 11 consulted no Chinese-language primary sources at all.
  2. Reliability grade inflation. 450 high, 152 medium, 1 low. A registry in which 0.2% of sources are low-reliability is not describing the information environment; it is describing a filtered environment. The filtering was correct — content farms were excluded rather than registered — but Phase 2 must not read the distribution as evidence that most sources are good.
  3. Selection toward the fetchable. 82.4% free, 55.7% public-page. Sources that blocked retrieval are under-registered relative to their value; several were registered only so the gap would be auditable (S-20-25 CARB, "Tier A source, zero coverage").
  4. Discovery-narrow sectors. Sectors 13–18 completed with 5–11 searches each instead of ~22 because the shared search budget was exhausted. Their Tier-A verification is sound; their source discovery is narrow, and their registry entries should be treated as a floor.
  5. Four missing types (patent, search_trends, social, community) and four near-empty ones (job_postings 1, open_source 1, conference 1, web_traffic 2) against a methodology that explicitly weights hiring and patents as leading indicators.
  6. One off-schema value (trade_org) indicates the enum is not validated on write. Fix at ingest.

Part 2 — The source-quality framework

6.5 A / B / C tiering and how to apply it

Tier A — primary. The record itself: SEC and other regulatory filings, company IR and press pages, government statistics agencies, central banks, standards bodies, patent offices, peer-reviewed papers, court documents. Tier B — reputable secondary. FT, WSJ, Reuters, Bloomberg, Nikkei, The Information, STAT News, trade press with named bylines, Big-4 and top-tier consultancy research. Tier C — everything else. Aggregators, vendor-sponsored market reports, SEO content farms, press-release wires, LLM-written blogs.

Three rules, all enforced rather than encouraged:

  1. Prefer A > B > C, always.
  2. Tier C may never be the sole support for a claim.
  3. Two outlets syndicating the same wire story are ONE source, not two.

Default tier by registry type

Registry type Default tier Promotion / demotion
regulatory_filing, government_data, standards_body, academic (peer-reviewed), company_ir, earnings, procurement, trade_data A Demote to B if the record is a summary page rather than the document; demote if undated
press_release A for facts about the issuer, marketing for claims about the market A vendor's claim about its own product is marketing and must carry the label
public_dataset A if the methodology is published (Visa Onchain, Berkeley Offsets, CVE); B otherwise Demote self-published crypto dashboards without disclosed adjustment criteria
trade_publication B with a named byline; C without Promote to A when it reproduces a primary document in full
newsletter B if the author is named and the analysis is original C if it aggregates without attribution
market_report B if the modeller and methodology are named and dated; C otherwise Industry-body statistical series (SEMI, WSTS, IFR) are B, and A for their own membership data
survey B, claims recorded as estimate C if panel size and composition are undisclosed
startup_db B, claims recorded as estimate Always name the tracker
app_store, web_traffic, developer_activity, search_trends B for declared/counted data, C for modelled estimates Never sole support
conference C Narrative signal only

Applying the tier at ingest

  • Every source needs a publication date. Undated web pages are Tier C at best and carry undated: true. Sector 01 found this is also a retrieval trap: "searching a current metric often surfaces a two-year-old article in the top results" — check the date of every result before use, and prefer the primary issuer's own current page.
  • Tier determines the evidence cap, and the cap is arithmetic. evidence_quality: 5 requires multiple Tier-A primaries; Tier-C-only caps it at 0–1; score = min(raw, 40 + 60E) where E = mean(evidence_quality, source_diversity)/5. A trend with weak evidence cannot exceed 40 no matter how loud it is. 25 of the 500 seed trends are capped by this rule.
  • source_diversity counts independent organisations, not links: 1 org = 1, 2 = 2, 3–4 = 3, 5–6 = 4, 7+ across at least two source types = 5.

6.6 Content-farm detection

Content farms dominate the search results for every sector in this programme. Sector 13 recorded them as "worse here, because humanoids are the most SEO-contested robotics topic"; sector 02 recorded that the entire public space for SaaS multiples and NRR benchmarks is content-farm-controlled; sector 09 found reshoring, warehouse-automation and industrial-AI-ROI statistics to be "almost entirely undated, unauthored pages citing unsourced percentages."

The signature — any three of these is a quarantine

Signal Test
Title template `/(Market (Report
No named author No author meta, no byline, or a byline that is a brand ("SEO DIGITAL PROS")
No date, or a rolling date Missing datePublished, or a date that changes on re-fetch
Unsourced forward figure A $XX billion by 20NN claim with no named methodology, sample or modeller
Monoculture The domain's entire content is market-size CAGR posts
Citation loop The only corroboration is other domains matching the same signature
Wire-press origin openpr.com and equivalents, which publish anything paid for
Statistic with no denominator "89% of AI pilots fail" attributed to a consultancy with no retrievable primary

Named and excluded in Phase 1

Recorded so the blocklist is auditable and seeded rather than rediscovered. From sector 06: lifesciencedaily.news, visionlifesciences.com, intuitionlabs.ai, peptidejournal.org, aimagicx.com, pdpspectra.com, humai.blog, presenc.ai, aimmediahouse.com, corstrate.com, nextaipress.com, deepceutix.com, biotechsign.com, hcranking.com, rxalmanac.com — several of which carried the only claims found on Isomorphic Labs' 2026 funding and on AI-drug-discovery market sizing, and were not laundered into the dossier. From sector 01: the "89% of enterprise agent pilots never reach deployment" figure, whose clearest attributing article carries the byline "SEO DIGITAL PROS" and no link, sample size or methodology — cited in the dossier only as an example of unsourced statistic propagation. From sector 11: search results for startup failure rates "dominated by Tier-C content farms recycling a decades-old '90% fail' claim."

Operational specification

  1. Deny-list the named domains and the openpr.com class. Blocking, not down-ranking.
  2. Score every new domain on the eight signals at first ingest; ≥3 → tier: C, quarantine: true.
  3. Quarantine is not deletion. Quarantined items remain retrievable as evidence of what is circulating, which is itself a signal — the divergence between what is claimed and what is verifiable is the fad diagnostic.
  4. Hard rule: a quarantined source can never be the sole support for a claim, and can never raise evidence_quality above 1.
  5. Track propagation. When a figure appears across ≥5 quarantined domains within 30 days with no primary, flag it as a propagating unsourced statistic and publish it as such. This is a product feature, not just a filter.

6.7 What independence actually means

source_diversity counts independent organisations. Four failure modes destroy apparent independence, and all four were observed in Phase 1:

  1. Wire syndication. The same story appears across dozens of domains. Sector 17 states the rule: "Syndication means the same story appears across many domains, which can create a false impression of independent corroboration — under this contract that is one source, not several."
  2. Shared upstream data. PitchBook underpins KPMG's venture figures (S-07-21), so citing both is not triangulation. EMARKETER and WARC observe the same advertiser panel universe (S-12-15) and must not be treated as independent confirmation of each other. Endpoints News is cited by other trackers as an underlying data source (S-06-15). Madison & Wall's concentration estimates were the most widely syndicated in 2026, so "apparent multi-source agreement often traces back to this one shop" (S-16-18).
  3. Republication chains. CnEVPost republishes both CPCA data and SNE Research battery tables; Silverado republishes Chinese customs data; the ISM PMI is most reliably obtained through its PR Newswire distribution. These are useful — often more machine-readable than the primary — but they are the primary, not a second source.
  4. Self-referential estimates. A vendor's market-size figure quoted back by the trade press that the vendor advertises in.

Dedupe algorithm for Phase 2:

for each incoming claim:
  normalise the numeric assertion (figure, unit, period, subject)
  cluster by (assertion, ±48h)
  within the cluster:
    resolve origin = earliest publication with an original byline,
                     or the named wire/agency, or the named data vendor
    all other members  -> republication, weight 0 for source_diversity
    members citing a shared upstream vendor -> collapse to that vendor
  source_diversity = count(distinct origin organisations)
                     across >= 2 distinct source_types for a score of 5

Record the republication chain rather than discarding it — chain length and speed is a useful measure of narrative velocity, and is exactly the input velocity is allowed to take.

6.8 Paywalled and licensed data

The rule, restated without qualification:

No scraping behind logins. No bypassing access controls. No circumvention of paywalls. No ToS violations. Prefer official APIs, RSS feeds, public pages and licensed data.

This is binding on Phase 2 engineering as it was on Phase 1 research, and it was observed: sector 01 read only the headline and dek of The Information's Anthropic gross-margin story and recorded the body as unread rather than obtaining it; sector 06 recorded STAT+ methodology as unverifiable rather than extracting it. An honest gap beats a plausible fabrication, and it also beats a ToS breach.

Four legitimate routes, in preference order:

  1. Free tier with attribution. Sufficient for direction in most cases. The free IFR press releases carry global installs, stock and density; the free SEMI release carries the global total; Green Street publishes monthly percentage changes; Circana headlines are released publicly via ESA and named public commentary.
  2. Licence the data. §6.9.
  3. Substitute. Berkeley's Voluntary Registry Offsets Database (S-20-09, free bulk) in place of the JS-rendered Verra registry; RBI payment-system indicators in place of NPCI; FFIEC CDR bulk in place of the FDIC QBP web pages; Justia (S-24-06) where CourtListener is blocked; on-chain supply in place of JS-rendered issuer dashboards.
  4. Record the gap. Publish the absence as a finding. Phase 1's most valuable negative results are of this kind: no frontier AI lab discloses a gross margin; no operator anywhere publishes robotaxi unit economics; no independent audited DPI benchmark exists for venture capital; no audited CPM index exists for any advertising channel.

6.9 Prioritised Phase 2 licensing budget

Ranked by: (a) no free substitute exists for the load-bearing series, (b) number of sectors served, (c) whether a scored trend currently depends on it, (d) monitoring cadence. List prices are not published for most of these and must be quoted; the cost band from the registry is given instead.

Tier 1 — buy first. No substitute; a scored trend depends on it.

# Source ID Band Sectors What the licence buys that free does not
1 IFR World Robotics (Industrial + Service) S-13-02 paid_high 13, 09 End-use, application and vendor-level breakdowns; the full national time series; the entire service-robot dataset. Free releases give only installs, stock and density. Annual series back to 1993. "Budget for this licence; there is no substitute." Sector 13 scores intelligence_demand 5 against data_availability 3 — maximum demand, worst information environment.
2 SEMI WWSEMS regional data S-03-10 freemium→paid 03, 09 The seven-region equipment table — the part that reveals China's equipment purchasing trajectory. The free release gives only the global total ($40.53bn, Q2 2026). Back to the 1990s.
3 WSTS detailed billings S-03-11 freemium→member 03 Monthly billings by product category and region, back to 1986. Member-restricted. "This is precisely why content-farm 'market size' pages fill the search results" — buying WSTS is buying the ability to refuse Tier C.
4 PitchBook / Morningstar S-07-21 enterprise 07, 11, 02, 18, 22 The private-markets database behind the $607bn / 567-fund evergreen universe. Note: it also underpins KPMG's venture figures — licensing it does not add an independent source, it identifies the shared one.
5 WARC S-16-17 enterprise 16, 12 Global ad-spend methodology and the effectiveness case database. The only source that surfaced independent academic evidence on AI creative performance rather than vendor lift claims. Sector 16's own instruction: "licensing budget should go to WARC first."
6 STAT+ (STAT News) S-06-14 paid_high 06 The M&A running totals used in sector 06 ($134bn YTD to 2026-06-22) sit behind it with their methodology. Per-seat licence; do not circumvent.
7 Circana US video game consumer spending S-15-11 paid_high 15 The authoritative US spending series back to the 1990s under NPD. Only headlines are free, via ESA.
8 Coresight store openings/closures databank S-12-17 paid_high 12, 18 The only weekly cadence in retail footprint data; line-item detail behind the paywall (the public preview gives ~7,900 closures / 5,500 openings for full-year 2026).

Tier 2 — conditional. Buy if the theme is tracked.

# Source ID Band Sectors Trigger for purchase
9 BloombergNEF (Energy Transition Investment Trends, NEO, battery price survey) S-05-19, S-10-13 paid_high 05, 20, 10 If energy-transition investment or battery cost curves are tracked. BNEF's data-centre demand trajectory (1,114 TWh by 2050) is materially more conservative than US utility forecasts — that divergence is itself worth tracking
10 Trepp CMBS loan-level S-07-19, S-18-11 enterprise 07, 18 If CRE credit is a tracked theme. Headline delinquency rates are free; loan-level and history are not
11 MSCI climate and carbon markets S-20-24 enterprise 20 If carbon prices or VCM volumes are needed. MSCI sells the weekly carbon credit price index and now gates First Street's research library. Together with Sylvera, BeZero and Abatable this is where VCM price and volume data actually lives, and none of it is free. Its own free pages are stale marketing (Net-Zero Tracker still dated 2024-08-31)
12 Comscore Box Office S-14-12 enterprise 14 If theatrical is tracked. The industry's actual currency; the public page returned no usable figures at all
13 The Information S-01-19 paid_high 01, 02 Currently the only outlet reporting frontier-lab gross-margin projections and cash-flow-positive timing
14 Green Street CPPI + NCREIF S-18-12, S-18-13 enterprise / paid_high 18, 07 Buy as a pair with a transacted index: REIT-implied (Green Street), appraised (NCREIF) and transacted. The spread between the three is more informative than any one of them
15 EMARKETER S-12-15 paid_high 12, 16 Only with WARC, and only with the independence caveat recorded on the claim
16 Crunchbase (data, not news) S-02-14 paid_high 01, 02, 04, 07, 08, 11, 22 Serves seven sectors; but Carta (free) covers the benchmark questions better and the Crunchbase/KPMG divergence must be carried regardless
17 DigiTimes / Nikkei Asia S-03-15 / S-03-17 paid_high / paid_low 03, 09 DigiTimes is the likely origin of much uncredited Taiwan supply-chain reporting — licensing it collapses several apparent sources into one, which is a quality purchase, not only an access one
18 Endpoints News, Secondaries Investor, GBTA BTI Data Cube, ConstructConnect, SNE Research, NCREIF, Aviation Week AWIN S-06-15, S-11-20, S-25-17, S-18-22, S-10-12 paid_high / enterprise 06, 11, 25, 18, 10, 08 Sector-specific; each named in its dossier's data-gaps section. SNE's free republished headlines are usually sufficient for detection

Tier 3 — defer

ISO/TC 299 standards text (S-13-20, paid_low, buy per standard on demand); conference programmes and access (S-16-24 — reliability: low, citation_value: low, buy only as a narrative instrument and never as evidence).

Not a licensing problem — an engineering problem

Four sources the brief and the dossiers flag as unobtainable are free, and were missed for retrieval reasons. Spending licence money on them would be an error:

Source ID Status Actual fix
Visa Onchain Analytics adjusted stablecoin volume S-21-12 free, daily, high reliability The dashboard is JS-rendered and returned methodology text but no figures. Fix: headless browser, or the underlying Allium data. This is the only source with a published methodology for separating genuine stablecoin payments from bot, bridge, MEV and exchange-internal traffic
CourtListener / RECAP S-16-22, S-17-14 free, official API, realtime Robots-disallowed through the Phase 1 session proxy. Fix: API token + declared user agent + rate limiting, per Free Law Project terms. Flagged in sector 16 as "HIGHEST-PRIORITY ADDITION FOR PHASE 2" — the Google remedies memorandum and proposed final judgment land on this docket, not in any press release
Carta Data Desk S-11-03 free, quarterly, high reliability Not paywalled at all. The gap is a publication gap — Carta had not published a Q2 2026 State of Private Markets at research date. Fix: monitor, don't buy. It remains the only free public source for down-round rates, stage-level round counts, AI vs non-AI valuation splits and fund-level DPI by vintage
CISA KEV / CVE Program S-04-02, S-04-22 free, machine-readable by design KEV JSON returned 403 to generic HTTP clients through some proxies; fix with a browser user-agent. For CVE use the cvelistV5 git repository, not the website

Highest-value free datasets still unexercised, which should be wired before any licence is signed: EU DSA Transparency Database (S-17-24), ESMA MiCA interim register (S-21-08), IMF PortWatch (S-09-23), OTEXA (S-23-23), SEC XBRL company-facts API (S-13-03, S-19-16), Berkeley Voluntary Registry Offsets Database (S-20-09), EIA API (S-05-01/S-20-14), Eurostat education and tourism (S-22-17, S-25-10), ENTSO-E Transparency Platform (S-05-11).

6.10 Access-blocker register

Engineering guidance. Every entry is a source the research needed, could not retrieve, and has a legal route to. None of the recommended routes involves circumventing an access control — every one is an alternative published endpoint, an official API, or a licence.

Class A — robots-disallowed

Source ID Blocker Recommended route
SEC EDGAR browse-edgar S-01-01, S-07-04, S-13-03, S-14-08, S-16-08, S-17-02, S-21-01, S-22-12, S-23-21, S-24-07, S-25-05 /cgi-bin/browse-edgar disallowed; efts.sec.gov 403s through proxies data.sec.gov/submissions/CIK##########.json to list filings → fetch the document under /Archives/. Use the XBRL company-facts API for figures. Declared User-Agent with contact, ≤10 req/s
congress.gov / CRS reports S-08-12 robots-disallowed govinfo.gov bulk data; Federal Register API for enacted instruments
NPCI UPI statistics S-07-23 npci.org.in robots-disallowed RBI monthly payment-system indicators or an official data partner
ISM Manufacturing PMI S-09-07 ismworld.org robots-restricted The PR Newswire distribution of the same release — machine-readable
Drewry World Container Index S-09-08 robots-restricted Licence, or secondary republication (Daily Cargo News). Fragile — treat as a licensing candidate
California Air Resources Board S-20-25 ww2.arb.ca.gov robots-disallowed; Tier A source, zero coverage California Regulatory Notice Register; CARB board-meeting PDFs by direct path; EPA/Federal Register for federal analogues
EUR-Lex search S-20-01 eur-lex.europa.eu/search.html robots-disallowed Direct legal-content URLs; find act numbers first via the EP Legislative Observatory (S-20-02)
OCC S-07-07, S-21-21 every occ.gov/occ.treas.gov path ROBOTS_DISALLOWED; robots.txt itself timed out Federal Register API for OCC rulemakings and bulletins; FFIEC for supervisory data
European Commission digital-strategy sub-paths S-17-06 several sub-paths robots-disallowed EUR-Lex + the Commission press-corner RSS
eSafety Commissioner (AU) S-17-05 robots.txt fetch timeout Direct PDF paths and the publications RSS
Swiss watch industry exports S-23-22 fhs.swiss robots failures Swiss Federal Office for Customs (Swiss-Impex)
OTEXA S-23-23 otexa.trade.gov robots connect timeout Bulk download from an unblocked egress; this is the single highest-value unexercised dataset in sector 23
Inditex, Kering S-23-24, S-23-25 hard site blocks; the sector's largest fast-fashion company contributed no 2026 data CNMV (Spain) and AMF / Euronext regulated-information filings — the statutory disclosure, not the corporate site
SpaceNews S-08-18 robots restricts some paths RSS feed only
CourtListener S-24-06 note robots-disallowed through the session proxy API token + declared UA (see §6.9); Justia (S-24-06) as interim substitute

Class B — JavaScript-rendered

Source ID Blocker Recommended route
Corporate IR index pages, sector-wide S-15-03, S-16-04, S-16-07, S-17-15, S-17-16, S-17-20, S-19-17, S-22-23, S-23-01, S-25-08, S-25-24, S-13-07, S-13-08, S-11-14, S-16-20 Index pages render client-side and return navigation shells Go straight to the individual press-release detail page or PDF — these fetch reliably. Confirmed independently in sectors 13, 16, 17, 19, 21, 22, 23, 25. gcs-web.com-hosted IR (S-25-06) serves detail pages cleanly. For financials, prefer EDGAR over the IR site entirely
Alphabet IR (abc.xyz) S-16-07 every automated retrieval failed; Alphabet's Q2 2026 advertising revenue is absent from sector 16 as a result EDGAR 10-Q via data.sec.gov/Archives/
Visa Onchain Analytics S-21-12 JS-rendered dashboard Headless browser, or Allium data
Tether / Circle transparency dashboards S-21-15 both JS-rendered, no figures returned Headless browser for live balances, or use on-chain supply instead
Verra registry S-20-21 registry.verra.org is a JS application Berkeley Voluntary Registry Offsets Database (S-20-09) bulk, or a registry API
Banco Central do Brasil Pix S-07-24 statistics page JS-rendered BCB open-data API endpoints
HKMA / SFC S-21-23 empty JS bodies or 404 on every attempt Headless browser or the RSS endpoint
AAR weekly rail traffic S-09-15 JS dashboard The weekly PDF, not the dashboard
NY Fed labour market for recent graduates S-22-01 interactive charts JS-rendered Locate the backing data files; fall back to the quarterly text
JNTO Japan inbound tourism S-25-20 jnto.go.jp 404s / client-side charts Japan Tourism Marketing Co. republication (registered as the working route)
Roblox IR S-15-03, S-17-15 JS-rendered; revenue, bookings and DAU not retrieved EDGAR 8-K and shareholder-letter exhibits

Class C — proxy 403 / 429 / rate-limited

Source ID Blocker Recommended route
CISA KEV JSON S-04-02 403 to generic HTTP clients through some proxies Browser user-agent; the JSON and CSV feeds are published for machine consumption
FDIC Quarterly Banking Profile S-07-01 QBP landing pages and PDF returned 403 FFIEC CDR bulk download or the FDIC BankFind API. /news/press-releases/ fetched fine
CBP trade releases S-12-10 cbp.gov 403 to automated fetch Newsroom RSS rather than page fetches
Monetary Authority of Singapore S-21-19 403 on every attempt RSS / subscription email
Federal Register /documents/search and /api/v1/ S-21-20 search robots-disallowed; API PROXY_REJECTED 403 govinfo bulk; direct document URLs (these work)
TSA daily checkpoint numbers S-25-15 tsa.gov 403 on every path, including with a browser UA Unblocked egress, or a mirrored dataset. Validate FAA/DOT (S-25-25) alongside
World Bank State and Trends of Carbon Pricing 2026 S-20-10 403 through the research proxy The Carbon Pricing Dashboard itself (instrument counts 47/40/34 were obtained)
UN Tourism Barometer S-25-22 403; news archive only pre-2025 National statistics offices; Eurostat tourism (S-25-10)
HESA (UK), OECD S-22-19, S-22-20 403 / every landing pattern 404 — no OECD figure appears in sector 22 HESA open data; OECD SDMX API rather than landing pages
ICE SEVIS / State Dept S-22-09 both 403 SEVIS by the Numbers archived PDFs; State Dept FOIA library
FSIS S-19-13 403 on news and proposed rules Federal Register (S-19-12) — the reliable substitute for FDA and FSIS pages that block
Artificial Analysis model pricing S-01-10 /models endpoint 429 repeatedly API or licensed feed. No current independent price-per-capability series was obtained at all
Assemblée nationale S-23-17 aggressive 429 Space requests; back off
USDA FAS agricultural trade sector 19 403 GATS / ESR bulk downloads

Class D — URL-pattern drift and archive 404s

Pattern Examples Rule
Deep-linked press releases 404 once archived Eli Lilly (S-06-10), Blackstone (S-11-14), The Trade Desk (S-16-20), Booking (S-25-08), Hilton (S-25-24), NIKE (S-23-07), Symbotic (S-13-07) Never store a constructed permalink. Resolve from the index or feed, store the resolved canonical URL and a content hash
Path-form traps FCA /news/press-releases 404s, /news works (S-07-13); Fed FEDS Notes year indexes use .html while most Fed pages use .htm (S-21-06); BIS work.htm / othp_list.htm 404 (S-07-10, S-21-07); SEC litigation-release path changed form (S-21-03) Maintain a per-publisher path map with a last-verified date; re-verify on every 404 rather than failing the source
Dated file URLs ESMA MiCA register CSV at a dated /sites/default/files/ path (S-21-08) Resolve the dated path from the landing page each run; never hard-code
Mixed fetchability within one agency FDA: FSMA 204 and front-of-package pages fetch, constituent-update and food-chemical-safety paths do not (S-19-11) Per-path health checks, not per-domain

General retrieval playbook (from the macro brief, confirmed by 25 sectors)

  1. SEC: data.sec.gov/submissions/CIK##########.json/Archives/. Declared UA, ≤10 req/s.
  2. Corporate IR: skip the index; fetch the release detail page or the PDF.
  3. Government statistics: direct file paths work (e.g. the Census e-commerce PDF at census.gov/retail/mrts/www/data/pdf/ec_current.pdf); navigation pages often do not.
  4. Federal Register is the universal substitute for any US agency page that blocks.
  5. Check the publication date of every result before using it. Date laundering in search results is systematic.
  6. Treat robots.txt as binding. Where a source is robots-disallowed, the route is an alternative endpoint, an official API, or a licence — never a workaround.

6.11 Monitoring design

Tier n Cadence Composition
tier1_daily 200 ≤24h 53 regulatory filings, 45 government data, 24 trade publications, 24 company IR, 11 earnings, 11 press releases, 7 pricing
tier2_weekly 246 ≤7d 44 government data, 36 trade publications, 33 regulatory/company IR, 22 market reports
tier3_monthly 157 ≤30d 41 government data, 33 market reports, 19 surveys, 14 academic

Engineering allocation follows from three counts:

  1. 386 sources (64%) can be polled through an API or RSS today — build this first; it is most of the value at the lowest cost. 317 records already carry a feed_url.
  2. 45 tier-1 sources have neither an API nor RSS. These are the highest-value manual surfaces in the registry and should get purpose-built fetchers with per-path health checks.
  3. 217 sources (36%) have neither. For these, a content-hash diff on a stable URL is the correct primitive: detect change, then parse — rather than parsing on every poll.

Cross-cutting requirements: deduplicate by endpoint and fan out by sector (§6.1); carry the legal_use_restrictions field into the fetcher as an executable policy (rate limits, declared user agents, attribution requirements, no-redistribution flags) rather than as documentation; and record every fetch failure against the source record so the access-blocker register in §6.10 stays current instead of becoming a snapshot of 2026-09-15.

6.12 The rule, one more time

No scraping behind logins. No bypassing access controls. No paywall circumvention. No ToS violations. Official APIs, RSS, public pages and licensed data only. Where that leaves a hole, the hole is published as a finding with the source registered so the gap is auditable — which is why S-20-25 (California Air Resources Board) sits in the registry as "Tier A source, zero coverage" rather than being quietly dropped.

Research provenance
Source artifact
01-frameworks/16-source-registry.md
Corpus date
15 September 2026
Prepared for this site
16 September 2026
Site publication
18 September 2026
Verification
Inherited; not fully rechecked