SASIGNAL ATLASCross-industry intelligence / Research desk
SIGNAL ATLAS / RESEARCH DESK

7 — Trend-Detection Methodology

Method · Cross-industry · Original Phase 1 research

7 — Trend-Detection Methodology

Phase 1 deliverable · research date 2026-09-15


7.1 The premise

Detection is not coverage. Every signal below is available to everyone; the ordinary trade press consumes most of them daily and still produces a stream of claims that cannot be audited. The difference between a feed and an intelligence product is entirely in the second half of this document: what you detect against.

Phase 1 produced the evidence for that claim. Across 500 scored trend records:

  • 104 records describe their own attention level as "extremely high" or "very high". Those 104 have a 27.9% overhyped rate against a 10% base rate — attention raises the odds that a trend is overhyped by roughly 2.8× — and a mean adoption score of 2.35 against a corpus mean of 2.93. Attention is not merely uninformative about substance; in this corpus it is negatively informative.
  • The 50 records classified overhyped carry the highest capital score of any cohort except current (2.44, above cooling's 1.77) and the lowest revenue (1.24) and customer_demand (1.68). Money present, demand absent. That two-number signature is the most reliable automated hype detector in the dataset.

This document specifies twenty signals, ranks them by lead time and precision, and then — the harder half — specifies eleven pathologies with a test for each, every one grounded in a named Phase 1 record rather than in the abstract. It closes with a pipeline an engineer can build.


7.2 Part 1 — The twenty detection signals

7.2.1 Signal reference

"Lead time" is relative to the point at which the trend becomes visible in revenue. "False-positive rate" is the proportion of times the signal fires and no trend follows, graded from the Phase 1 corpus and each sector's §7 record of prior false positives. "Registry supply" counts the 603 registered sources that can actually deliver it.

# Signal What it detects Lead FP rate How to measure Registry supply (of 603)
1 Repeated appearance across independent sources That a claim has more than one originating document — not that it is true 0 (corroboration test, not a detector) Very high if independence is not tested Count distinct originating documents after syndication collapse and ownership-graph collapse (§7.4.6) derived from all 603
2 Funding growth Investor belief; the size of the narrative's marketing budget +12 to +36m ahead of revenue High (~50%) Disclosed round value and count by category, rolling 4 quarters, median and count alongside the sum 11 startup_db + 2 vc_announcement = 13
3 Product launches Vendor commitment of engineering resource +6 to +18m Moderate–high Dated GA announcements; general availability only — previews, waitlists and "coming soon" are excluded 63 company_ir + 24 press_release + 3 app_store
4 Customer adoption Real deployed usage ~0 (confirming) Low Disclosed customer counts, seats, units, RPO, DAU — from the buyer's or platform's own accounts 30 earnings + 63 company_ir + 27 survey
5 Revenue growth Monetisation −3 to −6m (lagging) Very low Segment revenue in filings; reject non-GAAP-only presentations 92 regulatory_filing + 30 earnings
6 Search interest Consumer attention only 0 to +3m, consumer categories only Very high Normalised index, seasonally decomposed; never as a primary detector 0
7 Hiring Committed internal resource allocation — the cheapest honest signal there is +6 to +12m Low–moderate Role-title counts by employer; WARN filings for the reverse; job-postings indices 1 job_postings + S-02-04 (FRED/Indeed software postings index)
8 Patents R&D direction, 18 months before anyone talks about it +18 to +60m — the longest lead available Moderate (defensive and blocking filings) Filing counts by CPC class and assignee, accounting for the 18-month publication lag 0
9 Open-source activity Developer adoption ahead of procurement +6 to +18m Moderate Dependent-repo counts, package download volumes, distinct contributors. Stars are vanity and are excluded 1 open_source + 3 developer_activity
10 Regulatory attention What the state will permit, require or forbid +12 to +48m Low on existence, very high on timing Docket-stage tracking: proposed → enacted → in force → enforced, each dated separately 92 regulatory_filing + 130 government_data + 14 standards_body
11 Media coverage What other journalists think is interesting 0 to −6m (lagging) Extremely high Volume only, after syndication collapse; input to velocity and nothing else 73 trade_publication + 4 newsletter
12 Consumer behaviour Actual demand shift 0 to +6m Low if behavioural; very high if survey-based Transaction, traffic and unit data. Stated intent is an estimate, never a fact government_data retail series + 2 web_traffic + 3 app_store + 27 survey
13 Price changes Supply–demand imbalance, before volume moves +3 to +12m Low Spot and contract indices, ASPs, take rates, published price lists 10 pricing
14 New business formation Entrepreneurial entry into a category +12 to +24m Moderate Census BFS high-propensity applications; national registry equivalents S-11-05 (Census BFS)
15 M&A Incumbent conviction, priced 0 to +12m (confirming) Low on category validity, high on timing Deal counts and values by target category, from filings not from rumour regulatory_filing + press_release
16 Academic research The capability frontier +24 to +84m High — most published work goes nowhere commercially Publication counts by topic; citation velocity; replication status 25 academic
17 Conference activity Industry self-organisation around a category +6 to +18m High — programmes are substantially pay-to-play Track and session counts, exhibitor counts, keynote allocation, not attendance claims 1 conference
18 Government procurement Funded, contractual, non-reversible commitment +6 to +24m Very low — the highest-precision leading signal available Award notices, obligated dollars, contract vehicles, solicitation volume 3 procurement
19 Supply-chain shifts Physical reallocation that has already happened +3 to +18m Low Customs data, lead times, book-to-bill, export-licence grants and refusals 6 trade_data + government_data
20 Social / cultural discussion Attention among a self-selected, trivially manufacturable population 0 or negative Extremely high Only with bot filtering, account-age weighting and platform fraud adjustment 0

7.2.2 Ranking: lead time against precision

Lead time alone is the wrong ranking. Academic research has a five-year lead and a high false-positive rate; revenue growth has a negative lead and is almost never wrong. The usable ranking is two-dimensional.

High precision Low precision
Leading Tier 1 — build on these. Government procurement (18), supply-chain shifts (19), price changes (13), hiring (7), patents (8), regulatory attention (10, for existence) Tier 2 — hypothesis generators only. Funding growth (2), product launches (3), academic research (16), open-source activity (9), new business formation (14), conference activity (17), regulatory attention (10, for timing)
Coincident / lagging Tier 3 — confirmation. Customer adoption (4), revenue growth (5), M&A (15), consumer behaviour (12, behavioural only) Tier 4 — attention, not evidence. Media coverage (11), search interest (6), social discussion (20), consumer behaviour (12, survey-based), repeated appearance (1) if independence is untested

Operating rules that follow from the matrix:

  • A Tier-2 signal may open a record; it may never close one. Funding, launches and conference programmes create candidates. Promotion out of candidacy requires a Tier-1 or Tier-3 signal.
  • A trend supported only by Tier-4 signals is capped at evidence_quality ≤ 2, which under the §1 scoring model caps its composite at 64 and, if source_diversity is also low, at 40. This is arithmetic, not editorial preference.
  • Tier-1 signals are worth licensing budget; Tier-4 signals are not. The registry reflects the opposite priority today (§7.2.4).

7.2.3 Media coverage and social discussion are near-worthless as primary detectors

This needs saying without hedging, because it is the most expensive mistake in the category and the entire trend-intelligence market is built on the opposite assumption.

Media coverage measures editorial judgement about interest, on a one-to-six-month lag, with a syndication multiplier that makes one document look like thirty. The Phase 1 corpus contains four clean demonstrations:

  1. T-04-19 — the autonomous AI SOC. Coverage volume extremely high. A deliberate search for independent adoption or efficacy evidence in September 2026 returned exclusively vendor blogs, vendor-sponsored buyer's guides and consultancy content marketing. No neutral study, no regulator dataset, no peer-reviewed evaluation of autonomous triage accuracy in production. evidence_quality: 1, verification_status: unverified. Meanwhile the one independently measured operational metric that moved in 2026 moved backwards: median patching time rose from 32 to 43 days and KEV remediation fell from 38% to 26%.
  2. T-17-19 — creator-economy market sizing. The highest search volume in its sector and the lowest composite score in the entire 500-record corpus at 28.7. No multi-hundred-billion-dollar figure could be traced to a nameable methodology or modeller. The auditable numbers — YouTube's $100bn over four years (which covers creators, artists and media companies), Roblox's $1.5bn in 2025 — are an order of magnitude smaller.
  3. T-24-19 — "AI music is taking over streaming." Universal coverage, built on conflating Deezer's >50% of daily uploads with listening. The same disclosure says fully AI-generated music is 1–3% of total streams, and that up to 85% of the streams it does receive were fraudulent and were demonetised. Coverage volume was a perfect inverse indicator of the claim's accuracy.
  4. T-16-19 — "AI search is destroying demand capture." Used inside marketing organisations to argue for larger budgets while State Farm says on the record that budgets "have been essentially static for several years", and while the forecaster whose number is being cited (IAB, 12.3% US ad-spend growth) attributes the upgrade to the Winter Olympics and the FIFA World Cup.

Social and cultural discussion is worse, because it is cheap to manufacture. The decisive evidence is in sector 24: on Deezer, up to 85% of streams on fully AI-generated tracks were identified as fraudulent, against platform-wide stream fraud of 8%. When the engagement metric on a platform can be 85% bots for a given content class, no unadjusted social volume measure means anything. In sector 21, the SEC's Gotbit judgment (LR-26598, proposed final judgment filed 2026-07-28, with a parallel criminal guilty plea) establishes that manufacturing "the false impression of market interest" is a prosecuted practice, not a theoretical risk.

The rule the research contract already encodes, restated: search interest, funding volume, media volume and social volume are inputs to velocity and to nothing else. velocity is one of four momentum dimensions at 30% weight — 7.5% of the composite score. That number is the honest weight of attention in a calibrated system.

What media coverage is good for, and it is not nothing: it is a fast, cheap discovery layer for candidate generation, and it is the primary source for dated statements by named people, which is how opinion claims get captured with attribution. Both are real uses. Neither is detection.

7.2.4 The registry supplies the wrong signals — an audit

The 603 registered sources were chosen sector by sector, and in aggregate they are systematically misallocated against the ranking above.

Finding Evidence
Four signal types have zero registered sources patent 0, search_trends 0, social 0, community 0 — of 603
Four more have three or fewer job_postings 1, open_source 1, conference 1, developer_activity 3, procurement 3, app_store 3
The Tier-1 leading signals are the thinnest Patents 0, hiring 1, procurement 3, trade data 6, pricing 10 — 20 sources total for five of the six highest-precision leading signals
The lagging and confirming signals are the fattest government_data 130, regulatory_filing 92, trade_publication 73, company_ir 63, market_report 58
Only 2 of 500 trend records reference patents at all T-06-06, T-06-09 — both healthcare, where patent cliffs are unavoidable
Only 1 record references business formation T-11-18

Two of these absences are correct decisions: zero social and zero search_trends sources is consistent with §7.2.3 and with the contract's rule against attention as evidence. Zero patent sources is not. Patents are the longest-lead, moderate-precision signal available, they are free and bulk-downloadable from USPTO, EPO and WIPO, and their absence means the platform currently has no signal at all in the +18 to +60 month band. That is the single largest hole in the detection surface.

Phase 2 acquisition priority, in order: (1) USPTO/EPO/WIPO bulk patent data; (2) job-postings coverage beyond the single FRED/Indeed index — hiring is cheap, honest and currently served by one source; (3) procurement beyond US DoD/SAM.gov — the corpus is 29% b2g and has three procurement sources; (4) the EU DSA Transparency Database (S-17-24, already registered, daily bulk download, common schema across TikTok, Meta, Snap, Pinterest and X) as the only defensible way to observe platform-behaviour signals; (5) the access fixes named in the dossiers' §13 — TSA checkpoint throughput (HTTP 403), ww2.arb.ca.gov (robots-disallowed, which is why California SB 253/261 is entirely absent from the climate dossier), and the Berkeley voluntary-carbon-market workbook.

Access shape of the current registry, for the polling scheduler: 156 sources expose an API, 321 expose RSS, 62 offer bulk download, 12 require licensed access, 336 are public-page scrape-or-fetch. 497 are free, 78 freemium, 28 paid or enterprise. Monitoring priority: 200 tier1_daily, 246 tier2_weekly, 157 tier3_monthly.


7.3 Part 2 — What to detect against

Eleven pathologies. Each gets a mechanism, a tell, an implementable test, and a worked example from the Phase 1 research.

7.3.1 Early signals — recognising a true one

Mechanism. A genuine early signal is a divergence between a primary series and the prevailing narrative, observed in data that exists for a reason unrelated to the trend. It is not a small version of a trend; it is a hypothesis with evidence attached.

The tell. The signal appears in a series nobody was maintaining in order to make this point.

Test. Promote a candidate to emerging_signal only when (a) at least one primary series moves, (b) the mover is not the promoter, and (c) a falsifier is stated in the record's first_signals field with a date.

Worked example — T-21-01, stablecoin supply plateaus near $305bn. The narrative frame in 2026 was "the GENIUS Act unlocks growth". Five months of Federal Reserve, BIS and three independent tracker data showed no net expansion between April and September 2026 — and the plateau began before the regulatory regime took effect, which destroys the causal story in both directions. confidence: high, verification_status: triangulated, composite 77.3. Nobody was publishing that series to make a point about the GENIUS Act.

Counter-example of a correctly-marked weak signal — T-13-10, the humanoid component supply chain industrialising ahead of humanoid demand: real, dated, and scored 53.8 with stage: emerging_signal. The platform's job is to carry both and to say which is which.

7.3.2 False positives

Mechanism. The problem is real, the category forms, and the control turns out to be a feature rather than a product — or the timeline is off by a decade.

The tell. No one in the current marketing mentions the identical prior wave.

Test. Every sector dossier §7 must name that sector's own prior false positives, and the detector queries them as a lookup. A new candidate whose claim structure matches a listed prior false positive gets an automatic persistence penalty and an editorial flag.

Worked examples, from the dossiers' own records:

Sector Prior false positive 2026 candidate it should flag
04 SOAR (2017–21) — "eliminate Tier-1 analyst work through automated playbooks"; orchestration was easy, integrations and decision logic were not T-04-19 autonomous AI SOC — the identical promise, and the dossier notes SOAR "is almost never mentioned in current marketing"
03 450mm wafers — a decade and substantial capital, abandoned any industry-wide format transition requiring coordinated multi-party investment
05 Hydrogen for power (2020–24) — "announcement volume far exceeding docketed projects" SMRs; the dossier says to check them against exactly this pattern
11 Clean-tech 1.0 (2006–11) — most of ~$25B written off; returns depended on subsidies and commodity prices, not unit economics AI-infrastructure exposure with the same dependency structure
25 "Airlines will bypass the GDS with direct connect / NDC" (2012–present) — a fourteen-year-old imminent disruption T-25-19 agentic AI disintermediating the OTAs
25 "The internet will disintermediate travel agents" (1996–2005) — resolved as re-intermediation by Expedia and Booking every 2026 "AI will disintermediate X" claim
20 Clean Development Mechanism (2005–12) — "a UN-supervised methodology is sufficient proof of additionality" Article 6, which is reconstructing the same belief now
07 Marketplace/P2P lending (2013–16) — "banks will be disintermediated"; banks financed the disintermediators instead private credit in 2026

7.3.3 Coordinated promotion

Mechanism. A vendor cohort funds, launches and markets a category simultaneously. Analyst notes, buyer's guides, conference tracks and "independent" research all trace back to the cohort or to firms paid by it. The result reads like consensus.

The tell. High source count, low source independence, and every source has a commercial position in the outcome.

Test. Two automated measures on every candidate:

  • Promoter concentration = (citations originating from entities with a commercial interest in the claim) / (total citations). Flag above 0.5; refuse confidence: high above 0.7.
  • Commercial-interest flag on every source, computed from the source registry's bias_notes field plus an entity-relationship join: does this publisher sell a product whose market this claim sizes?

Worked example — T-04-19. Ten-plus named vendors marketing agentic SOC products simultaneously (Torq, 7AI, Qevlar AI, Conifers, Dropzone, Radiant, plus every platform incumbent), $300m+ of funding in twelve months, and a promoter concentration of effectively 1.0: every locatable source was vendor-produced or vendor-sponsored, including the consultancy content marketing. The two Tier-C sources in the record are aggregator posts assembling vendor data.

Worked example — T-16-19. Subtler and more instructive: the promoters are the buyers' own agencies. Digiday documents brands in insurance and automotive using AI-search anxiety to argue internally for expanded media budgets. The party generating the signal is the party that benefits from the budget it justifies.

7.3.4 VC hype

Mechanism. Capital is a leading indicator of belief, not of adoption. A large round is, functionally, the marketing budget for a narrative — and in 2026 the narrative is frequently the product.

The tell, quantified from the corpus. The overhyped cohort carries capital 2.44 — higher than the cooling cohort's 1.77 — against revenue 1.24 and customer_demand 1.68. Money in, demand out. Compute capital − mean(revenue, customer_demand) on every record; a value ≥ 2 is the hype signature.

Test.

  1. Does the funding announcement disclose revenue? T-24-19's evidence records that Suno's $5.4bn valuation announcement (2026-06-03, $400m raised) "contains no revenue, ARR or subscriber figures" — captured as claim_type: signal, not as fact.
  2. Capital-to-verifiable-revenue ratio. T-13-19: roughly $100bn+ of aggregate private and public valuation against $1.8m of disclosed 2025 revenue at Agility Robotics with a $140m operating loss, Tesla at zero Optimus units performing useful work as of January 2026 against a promised 10,000 in 2025, and less than 10% of Unitree's 2025 revenue from industrial applications. The dossier calls this "the largest attention-to-substance gap in this research programme," and every number in it comes from a prospectus, an S-4 or the sector's own trade body — not from a sceptic.
  3. Who is the capital? Sector 11 records corporate VC at 87.9% of US AI venture deal value. Strategic capital withdraws on a strategy review, not on a fund cycle — a completely different persistence profile that the aggregate "venture funding" number hides.

7.3.5 Media echo chambers

Mechanism. N outlets, one originating document. Each restatement increments the naive corroboration count without adding information, and an advocacy body restating a vendor statistic converts a commercial claim into apparent independent support.

The tell. All citations reduce to one primary document when you follow them back.

Test — originating-document resolution. For every claim, resolve the citation chain to the earliest primary document and count distinct originating documents, never distinct URLs. The research contract already states the rule: two outlets syndicating the same wire story are ONE source.

Worked example — the Deezer loop (T-24-04, T-24-19). Deezer publishes the AI-upload statistic (2026-07-21). IFPI reproduces the Deezer figures in its EU policy advocacy (2026-09-01) — correctly captured in T-24-04 as claim_type: opinion with the note that IFPI "amplifies the supply-side statistic in a context where a demand-side threat is being argued". Trade press then cites both, and the naive independent-source count reads 2+. The true count is 1.

Worked example — the YouTube $100bn loop (T-17-04, T-17-19). CNBC reported "$100 billion over four years" on 2025-09-16; Neal Mohan's CEO letter reported the same $100bn over four years on 2026-01-21, sixteen months later. Same rounded figure, two dates, and the earlier Forbes Australia article reporting $70bn over three years (2024-02-07) still surfaces prominently in September 2026 search results without recency signalling. This is date laundering: a stale figure re-entering circulation as current because search ranking does not encode recency.

Two hard rules that follow:

  • Every retrieved document must have its publication date extracted and validated before use. Undated pages are Tier C and flagged undated: true. The macro brief lists date laundering as a confirmed cross-sector trap.
  • Rounded cumulative disclosures cannot be differenced. YouTube's "$100bn over four years" appeared at the same $100bn a year earlier, so no annual growth series can be derived from it. The pipeline must refuse to compute a delta between two roundings.

7.3.6 Fads

Mechanism. Attention rises without a corresponding rise in adoption, revenue or capability. A single dominant promoter class; no independent demand.

The tell — the divergence signature. High velocity, low adoption, low persistence, low evidence_quality.

Test. velocity ≥ 4 AND adoption ≤ 1 flags 26 of 500 records, of which 5 are already overhyped and 14 are emerging_signal — exactly the cohort that needs a persistence check. Combine with the contract's enforced floor: a trend under twelve months old cannot score above 2 on persistence. That single anchor is the anti-fad mechanism and it is arithmetic, not judgement.

Worked example — the metaverse as general-purpose computing platform. Meta's Reality Labs lost $8,647m in H1 2026 on $833m of revenue, roughly 10:1, and no credible market model carries a separate metaverse revenue line. Classic signature.

Worked example in-corpus — T-23-19, GLP-1 drugs reshaping apparel sizing and beauty demand. evidence_quality: 1, source_diversity: 1, verification_status: unverified, composite 29.3 — the second-lowest in the corpus. Plausible mechanism, zero measurement.

7.3.7 Seasonal patterns

Mechanism. A calendar effect is read as a directional change. This is the most common innocent error in the corpus and the easiest to automate away.

The tell. The comparison window is shorter than a year, or the series is not-seasonally-adjusted and is being compared period-over-period.

Test. Three rules, all mechanical:

  1. No direction is asserted from fewer than 13 months of series. Year-over-year, never week-over-week or month-over-month, unless the publisher provides a seasonally adjusted variant — and if it does, use it and say so.
  2. Check whether a moveable holiday sits in the comparison window.
  3. Check whether an event calendar explains the move before attributing it to a structural cause.

Worked example — rule 2 (25 §13). US hotel RevPAR printed +4.4% in the week to 2026-08-22 and +16.1% in the week to 2026-09-05. Pure Labor Day calendar distortion. Single-week STR prints are unusable without calendar adjustment, and secondary coverage routinely quotes whichever one supports the story.

Worked example — rule 3 (T-16-19 contradictions array). The IAB raised its 2026 US ad spend growth forecast to 12.3% from 9.5%. Marketers quoted in Digiday attributed the upgrade to "adaptation to changing search habits." The forecaster's own stated driver was a stronger-than-expected H1 driven by the Winter Olympics and the FIFA World Cup. A sporting calendar was retrofitted into a strategic AI narrative.

Worked example — rule 1. Sector 25's US domestic demand finding is stated as seasonally adjusted US passengers at 80.1m in June 2026, 3.7% below the June 2024 peak of 83.2m, with three consecutive months of declining domestic RPK. Three consecutive months plus a two-year comparison is the minimum defensible form.

7.3.8 Market bubbles

Mechanism. Price detaches from the cash flow it claims to represent. The detection problem is that the price itself is widely reported as evidence for the trend.

The tell. Valuation moves violently while operations do not move at all.

Test. Track the realisation ratio — cash actually returned or revenue actually booked, against capital deployed or valuation ascribed — and report medians and counts alongside every sum.

Worked examples:

  • T-13-16 — Unitree. IPO subscription oversubscribed 2,760.67×; +460% on debut; peak valuation RMB 445bn ($66bn); by 2026-09-09 down 53% from the first-day high and 39% from the first-day close, erasing roughly $35bnwhile nothing changed operationally. HSBC had warned pre-IPO that "the surge in shipments for robot makers could be illusionary." This is now the public comparable against which private humanoid marks — including Figure AI's $39bn September 2025 valuation — must be tested.
  • SpaceX / SPCX. The largest IPO in history ($75bn raised at a $1.5T valuation, 2026-06-12, day-one close +19%) is −30.3% from debut, and its first IR disclosure shows the Space segment at $962m of $7.814bn Q2 revenue and the only loss-making core segment (−$542m operating), with Starlink ARPU down from $85 to $66.
  • T-11-01 — the paper/cash divergence. 2021-vintage DPI 0.05x; net LP cash flow since 2022 −$202B; four companies = 93.5% of 2026 exit value; seed→Series A graduation down from 55%+ to 16% for the 2024 cohort. Aggregate dollars look euphoric; the median company's experience is the opposite. Never cite headline VC totals without this.

7.3.9 Disinformation

Mechanism. Signals are deliberately fabricated. Distinct from promotion: the underlying metric is falsified, not merely emphasised.

The tell. The metric is one a participant can manufacture at low cost and has a direct financial reason to manufacture.

Test.

  1. Refuse any engagement or volume metric the platform has not fraud-adjusted, and record the adjustment.
  2. Check the promoter's exposure. Who profits if this number is believed?
  3. Check the enforcement record. If a regulator has prosecuted manufacture of this specific signal in this sector, the prior is not neutral.

Worked examples:

  • T-24-04. Deezer identified up to 85% of streams on fully AI-generated tracks as fraudulent in 2025 and demonetised them, against platform-wide stream fraud of 8%. The commercial motive is royalty extraction. Any "AI music listening share" computed from raw streams is measuring bots.
  • Sector 21, SEC v. Gotbit (LR-26598, 2026-07-28, with a parallel criminal guilty plea). Wash trading to "create the false impression of market interest" is a prosecuted practice in crypto. Sector 21's operative rule is the correct general one: a claim counts only if it appears in an audited filing, a regulator's document, a fund's daily NAV, or posted margin.
  • Sector 04, "Q-Day is imminent." Published estimates for a cryptographically relevant quantum computer span the early 2030s to the mid-2040s. The near-term dates circulating in 2026 come predominantly from cryptocurrency-community figures with direct exposure to quantum-vulnerable signature schemes — not from quantum hardware groups. Not fraud, but interested forecasting presented as technical assessment, and the test that catches it is the same one.

7.3.10 Survivorship bias

Mechanism. The measured sample is the set of things that survived long enough to be measured. Every conclusion drawn from it is conditioned on survival, silently.

The tell. Ask what the denominator excludes, and the answer is "we don't collect that."

Test. For every rate or average, require the denominator's construction. Where the denominator is self-selected, cap evidence_quality at 3 regardless of source tier and state the selection mechanism in the record.

Worked examples:

  • Venture performance (T-11-01, sector 11 §13). Every DPI figure comes from PitchBook, Carta, Cambridge Associates or McKinsey — "each with a different, self-selected, survivorship-affected sample and no external audit. No regulator publishes fund-level performance." The direction triangulates across four organisations; the level is not verifiable to Tier-A standard. Sector 11 calls this its largest gap.
  • Startup failure. Sector 11 could find no reliable shutdown statistic: search results were dominated by Tier-C content farms recycling a decades-old "90% fail" claim (excluded entirely), Carta's dissolution series was unpublished for 2026, Census BFS measures formation not venture-backed failure, and BLS Business Employment Dynamics lags roughly two quarters. The failure rate beneath the AI mega-rounds is the least-evidenced part of that dossier, and it is the denominator for every success claim in it.
  • The invisible majority. Sub-$50M funds are 67.7% of fund closings and 4% of dollars. Any dollar-weighted analysis erases the segment where most first cheques originate.
  • Attention-side survivorship (sector 20). The durable-CDR delivered-tonnes leaderboard is entirely Global South biological routes — Exomad Green (Bolivia, 416,703 t), Varaha (India, 185,841 t), Carboneers, Aperam BioEnergia, O.C.O. Technology. The most successful carbon-removal company in the world by the only metric that counts is a Bolivian biochar producer, and it is invisible because media gravity follows contracted volume, a different leaderboard entirely. 1.68 Mt delivered against 49.47 Mt contracted.

7.3.11 Data gaps

Mechanism. Absence of data is read as absence of the phenomenon, or — worse — is silently filled with the nearest available number of a different construction.

The tell. The dossier has nothing to say about something important, and does not say so.

Test. data_gaps is a mandatory, non-empty array. The research contract states plainly that a dossier with an empty data_gaps array will be rejected as unserious. Every gap records what was sought, why it failed (paywall, robots.txt, 403, JS-rendered, never published), and what the consequence is for the conclusions.

Worked examples — gaps that are themselves findings:

Gap Consequence
TSA daily checkpoint throughput returns HTTP 403 on every path attempted The highest-frequency independent measure of US air-travel demand is unavailable; sector 25's US demand finding rests on monthly, lagged BTS and IATA data. Highest-priority Phase 2 access fix
ww2.arb.ca.gov is robots-disallowed California SB 253 / SB 261 — the most significant mandatory US corporate climate disclosure requirement — is entirely absent from the climate dossier
No VCM transaction volume, market value or average price for 2025 or 2026 — Ecosystem Marketplace carries no figures, Berkeley publishes only inside Excel, registry.verra.org is a JS app, MSCI is paywalled The voluntary carbon market's size appears nowhere in sector 20, and none should be inferred from it
No frontier AI lab discloses a gross margin Every AI unit-economics claim in the corpus is modelled, not measured
No operator anywhere publishes robotaxi unit economics Same
No independent, regulator-published, fund-level venture performance dataset exists §7.3.10
Only 24% of H1 2026 US breach notices disclosed an attack vector — the lowest ITRC has recorded The public evidence base for cybersecurity is degrading, handing vector data to vendors as the only remaining source
EA completed its take-private on 2026-08-04 A top-five publisher's quarterly reporting disappeared. Data availability is itself a trend, and it is moving the wrong way

And the statistical traps that fill gaps incorrectly, confirmed across multiple sectors and required as pipeline assertions:

  1. US Census reports data-centre construction inside the "Office" category — which is why Census office construction is +21.3% y/y in the middle of an office-distress narrative. Any nonresidential construction series must be decomposed before use.
  2. Core CPI 2.4% (Aug 2026) against core PCE ~3.3% — an unusually wide gap in the unusual direction. Never say "inflation"; say which index.
  3. IATA against itself: its monthly press release and its Air Passenger Market Analysis report different figures for the same region and month — North America −2.3% vs −1.2%, Asia-Pacific −0.7% vs +1.0%, opposite signs. Different aggregation bases. The two products must never be mixed in one series, and secondary coverage does exactly that.
  4. Definitional spreads presented as single numbers: tokenised RWA "represented asset value" $361.91bn against "distributed asset value" $38.89bn — a ninefold spread that commentary routinely collapses.
  5. Aggregate versus median. True in every sector. Report medians and counts alongside sums, always.

7.4 The detection pipeline

Ten stages. Written to be implemented.

7.4.0 Source registry and polling scheduler

The registry is the system of record for what may be polled and how. Each of the 603 sources carries access_method, update_frequency, monitoring_priority, cost, legal_use_restrictions, reliability, bias_notes, data_quality and citation_value.

schedule(source):
    cadence = {tier1_daily: 24h, tier2_weekly: 7d, tier3_monthly: 30d}[source.monitoring_priority]
    jitter  = uniform(0, 0.1 * cadence)
    method  = source.access_method    # api > rss > bulk_download > public_page
    assert source.legal_use_restrictions allows programmatic retrieval
    assert robots_txt_allows(source.url)     # hard gate: no login-walled, no paywall bypass

Ordering is deliberate: API before RSS before bulk download before page fetch. Of 603 sources, 156 expose an API, 321 RSS, 62 bulk download. The 336 public-page sources are the fragile tail and need per-source fetch recipes — the retrieval playbook in the macro brief is the starting set (data.sec.gov/submissions/CIK##########.json rather than browse-edgar; individual IR press-release URLs rather than JS-rendered index pages; direct government file paths rather than navigation pages).

Failures are recorded, not swallowed: a source that 403s or goes JS-only becomes a data_gap row, which is what turns an access failure into a Phase 2 licensing decision instead of a silent hole.

7.4.1 Fetch and normalise

doc = fetch(source)
doc.content_hash      = sha256(normalised_text(doc))
doc.canonical_url     = strip_tracking_params(resolve_redirects(doc.url))
doc.published_at      = extract_date(doc)   # <time>, JSON-LD, OpenGraph, byline, HTTP headers
if doc.published_at is None:
    doc.undated = True; doc.tier = "C"      # contract §0.9
if doc.published_at older than 180 days and the claim is presented as current:
    flag DATE_LAUNDERING                    # §7.3.5

Date extraction is not a nicety. It is the defence against the single most common retrieval error in the Phase 1 research.

7.4.2 Deduplication and syndication collapse

near_dupes = minhash_cluster(docs, jaccard_threshold=0.82)
for cluster in near_dupes:
    origin = earliest_published(cluster)
    for d in cluster: d.syndicated_from = origin.id

Additionally resolve explicit citation chains: when document B's load-bearing figure is attributed to document A, set B.derives_from = A. Both edges feed the independence graph.

7.4.3 Entity and claim extraction

entities = ner(doc) |> resolve_against(entity_registry)   # global IDs, not strings (§3.3.5)
claims   = extract_numeric_claims(doc)
for c in claims:
    c.value, c.unit, c.currency, c.period, c.basis = parse(c)
    c.claim_type = classify(c)        # fact | signal | estimate | forecast | opinion | marketing
    c.subject_entity = link(c, entities)

claim_type classification is the highest-value model in the pipeline. The rule that makes it load-bearing: a marketing claim may never be rendered without its label. In the seed corpus, 46 of 2,580 evidence items (1.8%) are marketing — including Deezer's own "99.8% accuracy" statement about its detector, correctly labelled inside T-24-04 and T-24-11.

7.4.4 Definitional-trap detection

Run before clustering, because a trap that survives to this point becomes a trend.

TRAPS = [
  ("supply vs demand",       {uploads, listings, registrations} vs {streams, sales, active users}),
  ("contracted vs delivered",{contracted, announced, committed} vs {delivered, docketed, in service}),
  ("represented vs distributed", {represented AUM, notional} vs {distributed, realised}),
  ("gross vs net",           {gross bookings, GMV} vs {net revenue}),
  ("statutory vs adjusted",  {net profit} vs {adjusted, ex-non-recurring}),
  ("cumulative vs annual",   {"over N years", "to date"} vs {annual}),
  ("index mismatch",         {CPI} vs {PCE}),
  ("SA vs NSA",              seasonally adjusted vs not),
]
if two claims in a candidate differ only by basis: emit CONTRADICTION, do not reconcile

Every one of these is a real Phase 1 instance: uploads vs streams (T-24-04, >50% vs 1–3%); contracted vs delivered (sector 20, 49.47 Mt vs 1.68 Mt); represented vs distributed (sector 21, $361.91bn vs $38.89bn, 9×); statutory vs adjusted (T-13-16 — Fortune reported Unitree's adjusted RMB 590.75m as "RMB 600m profit" against RMB 278.21m statutory, roughly doubling apparent profit); cumulative vs annual (YouTube's $100bn); CPI vs PCE (2.4% vs 3.3%).

The output is a contradictions[] entry, not a resolution. The corpus carries 201 contradictions across 500 records, and the mission document is explicit that this is a feature: "where the honest answer is 'credible sources disagree and here is why,' that is the answer."

7.4.5 Clustering into candidate trends

candidate = cluster(claims) by:
      embedding_similarity(claim_text) > θ₁
  AND entity_overlap(subject_entities) > θ₂
  AND |published_at difference| < 90 days
    → assign industry_id by routing rules (§3.3.1: market of the buyer)
    → assign trend_type family (§3.2) by the evidence class present
    → assign to trend_cluster if it matches an existing cross-sector cluster (§3.5.1)

The cross-sector cluster step is the one the seed database is missing entirely: 0 of 1,535 related_trends links cross an industry boundary, while 33 records across 9 sectors turn on data-centre demand and 30 across 15 sectors on app-store commission economics.

7.4.6 Corroboration counting and the independence test

This is the stage that distinguishes the product. Naive corroboration counts URLs; independent corroboration counts causally independent originating organisations.

def independent_sources(candidate):
    nodes = {c.source for c in candidate.claims}

    # 1. collapse syndication and citation chains  (§7.4.2)
    nodes = {root_of(n) for n in nodes}

    # 2. collapse common ownership / control
    nodes = {ultimate_parent(n) for n in nodes}          # publisher → group → owner

    # 3. collapse commercial dependence
    #    a source funded by, sponsored by, or reselling the promoter is NOT independent
    nodes = {n for n in nodes if not commercially_dependent(n, candidate.promoters)}

    # 4. collapse shared upstream data
    #    two "independent" trackers reading one registry are one source
    nodes = collapse_by_upstream_dataset(nodes)

    # 5. count source TYPES, not just organisations
    types = {n.source_type for n in nodes}
    return len(nodes), len(types)

orgs, types = independent_sources(candidate)
source_diversity = 1 if orgs==1 else 2 if orgs==2 else 3 if orgs<=4 else 4 if orgs<=6 else (5 if types>=2 else 4)
verification_status = "triangulated" if orgs >= 2 else ("single_source" if orgs == 1 else "unverified")

Steps 3 and 4 are the ones ordinary aggregation omits, and they are where the Phase 1 examples bite.

Step 3, worked — Deezer. Deezer is simultaneously (a) the sole source of the AI-upload statistic, (b) the vendor of the AI-detection product that the statistic justifies — it "began licensing its AI music detector, deploying it in 27 languages and monetising it through agreements with Hungarian and Dutch collecting societies" (T-24-11) — and (c) the party whose own fraud-adjustment (85%) is required to interpret its own demand-side number. A source that sells the remedy for the problem its statistic establishes is one interested source, and T-24-19 is correctly marked verification_status: single_source with evidence_quality: 2 despite two Tier-A citations, because the second (IFPI) is reproducing the first.

Step 4, worked — stablecoin trackers. Sector 21 records three "independent" stablecoin trackers disagreeing by 4% on the sector's single most-cited number. They are not three sources; they are three readings of the same on-chain state, and their agreement carries far less information than three genuinely different measurement methods would.

The promoter-concentration and commercial-interest computations of §7.3.3 run here, off the same graph.

7.4.7 Signal scoring

Map the twenty signals onto the fifteen dimensions, then apply the cap.

velocity   ← signals 2,3,6,11,20     # attention-heavy: HARD CAP AT 3 if the only inputs are 6,11,20
adoption   ← signals 4,9,12          # no measurable users ⇒ 0–1, no exceptions
capital    ← signals 2,15,18
revenue    ← signal 5
breadth    ← cross-sector cluster membership (§7.4.5)
depth      ← editorial
geographic_spread ← geography facet cardinality and tier (§3.4)
customer_demand   ← signals 4,12,18  # vendor-push-only ⇒ 1
persistence       ← age of first observed signal; < 12 months ⇒ max 2
technical_maturity← signals 8,9,16,3
strategic_importance, regulatory_impact, social_impact ← editorial, signal 10 informs the second
evidence_quality  ← source tier mix
source_diversity  ← independent_sources() from §7.4.6

raw   = 100 * (0.30*M + 0.25*R + 0.25*D + 0.20*C)
E     = mean(evidence_quality, source_diversity)/5
score = min(raw, 40 + 60*E)

The cap bound 25 of 500 seed records. It is the mechanism, not the intention, by which a loud trend with thin evidence cannot outrank a quiet one with good evidence.

7.4.8 Pathology detectors

Each of §7.3's eleven runs as an automated flag. Flags do not block publication; they route to the editorial gate and are surfaced on the record.

Flag Fires when
PROMOTER_CONCENTRATION interested-source share > 0.5
ECHO_CHAMBER orgs after collapse < 0.4 × raw citation count
VC_HYPE capital − mean(revenue, customer_demand) ≥ 2
FAD_SHAPE velocity ≥ 4 AND adoption ≤ 1
SEASONAL_RISK comparison window < 13 months, or a moveable holiday in the window
BUBBLE_RISK valuation change > 30% with no change in disclosed operating metrics
FRAUD_EXPOSED_METRIC the metric is platform-engagement and no fraud adjustment is published
SURVIVORSHIP the denominator is self-selected or vendor-constructed
HISTORICAL_ANALOGUE claim structure matches a listed prior false positive in that sector's §7
DEFINITIONAL_TRAP §7.4.4 fired
DATA_GAP a registered source for a required signal failed or does not exist

7.4.9 The editorial gate

Automation produces candidates, scores and flags. It does not publish. Five things require a human and are not delegable:

  1. Claim-type adjudication on every marketing and opinion item. The rule that unlabelled vendor claims never render is the system's single most important editorial control, because that is the primary vector by which hype enters research.
  2. The independence judgement in step 3 of §7.4.6. Whether a source is commercially dependent on a promoter is a factual question the graph can propose and a human must confirm. The Deezer case is not detectable from metadata alone.
  3. Contradiction framing. The pipeline detects that two figures differ; a human writes likely_reason — methodology, scope, timing or definition. All 201 seed contradictions carry one.
  4. classification assignment, specifically overhyped. Calling something overhyped is an accusation and must be defensible from the record's own evidence. Note the discipline in the corpus: T-13-19 is overhyped with confidence: high and evidence_quality: 5, because every falsifying number came from a prospectus, an S-4 or the trade body — not from a sceptic.
  5. The falsifier. No trend publishes without a stated, dated indicator that would change the conclusion. Sector 04's example of the form: the soft cyber-insurance market reverses on a correlated-loss event; early indicator, any carrier publicly restating cyber reserves.

Hard publication blocks: Tier-C-only evidence; empty data_gaps on a dossier; a marketing claim rendered without its label; a GLOBAL geography tag with geographic_spread < 4; a record asserting direction from under 13 months of series.

7.4.10 Publication, versioning and re-verification

Every record carries last_verified and a re-verification SLA keyed to time_horizon: immediate_0_12m monthly, near_1_3y quarterly, medium_3_7y semi-annually, long_7y_plus annually. Re-verification re-runs §7.4.1–7.4.7 against the same sources and diffs the scores, so that a trend's trajectory through the corpus is itself a series.

The falsifier register is the product's memory: when a stated falsifier fires, the record is re-opened and the outcome recorded whether or not it was flattering. That register — not the trend count — is what makes the platform's calibration auditable over time, and it is the only thing that distinguishes a forecast from a prediction.


Companion documents: 10-mission-and-definitions.md (§1, trend vs signal vs fad, the claim-type ladder), 12-taxonomy.md (§3, the facets this pipeline writes), 11-industry-ranking.md (§2), 01-research-contract.md (the source-tier rules and the scoring rubric this methodology enforces).

Research provenance
Source artifact
01-frameworks/13-detection-methodology.md
Corpus date
15 September 2026
Prepared for this site
16 September 2026
Site publication
18 September 2026
Verification
Inherited; not fully rechecked