SASIGNAL ATLASCross-industry intelligence / Research desk
SIGNAL ATLAS / RESEARCH DESK

44 — Case Studies

Analysis · Cross-industry · Original Phase 1 research

44 — Case Studies

Eight worked examples of the methodology doing something a news aggregator could not Research date 2026-09-15


0. What these are for

Each case below has the same structure: what was claimed · what the evidence showed · how the methodology caught it · which scoring dimensions moved · the generalisable lesson.

The selection criterion is narrow and deliberate. Every case is one where an aggregator — or any process that reports what credible sources say — would have produced a confidently wrong answer, because the sources were credible and the reporting was accurate. The error in each case is not in the facts but in the relation between facts: a definition, a denominator, a date, a legal basis, or an absence.

The scoring model is referenced throughout. Its relevant arithmetic: fifteen dimensions scored 0–5, a composite normalised to 0–100, an evidence factor E = mean(evidence_quality, source_diversity)/5, and a hard cap at 40 + 60E. Attention is an input to velocity only, and velocity is 7.5% of the composite.


Case 1 — Crunchbase vs KPMG: a $50.4bn disagreement that is not an error

The canonical worked example for the source-quality framework

What was claimed. Global venture funding for H1 2026. Both figures are published by reputable trackers and both circulate as fact.

Metric Crunchbase KPMG Venture Pulse Gap
H1 2026 global VC $510bn $560.4bn $50.4bn
Q2 2026 global VC $205bn $227.4bn $22.4bn
Q2 2026 deal count "5,000+" 8,440 ~3,400 deals

What the evidence showed. Neither is wrong. The divergence is methodological: deal inclusion rules, treatment of corporate and private-equity rounds, and announced-versus- closed dating. A $50.4bn gap is larger than the entire annual venture market of most countries, and the deal-count gap is proportionally far larger than the dollar gap — which is itself the tell, because it says the two trackers disagree most about small deals, the population that matters for T-11-18 and T-11-03.

Crucially, both agree on direction: a record half-year with a collapsing deal count. KPMG's 8,440 deals in Q2 2026 is the lowest count since Q3 2017.

How the methodology caught it. Three mechanisms, in order:

  1. The contradiction is recorded as a contradiction, not resolved. The database holds 201 recorded contradictions with figure, source_a, value_a, source_b, value_b and likely_reason. The ordinary flow of trade press resolves this by silently picking one number; the platform is forbidden from doing so.
  2. The estimate claim type requires a named modeller. Neither figure may be rendered as fact. Both are modelled counts of a population neither tracker observes completely.
  3. The count/sum rule. Because both series were recorded, the divergence in the ratio became visible, which is what made the aggregate/median finding (T-11-01, T-11-02, T-11-03) available at all.

Which scoring dimensions moved. On T-11-02 (four-layer capital concentration): the concentration facts are triangulated across Crunchbase, KPMG, PitchBook and Carta, so evidence_quality and source_diversity both score 5, E = 1.00, cap = 100, composite 82.3 uncapped. But revenue scores 3 and customer_demand 3 — the model refuses to let a capital-flow record score as though it were a revenue record. On T-11-03 (seed→A graduation collapse), revenue scores 1 and regulatory_impact 1, holding the composite to 72.8 despite five-point adoption, breadth, depth and technical maturity.

Generalisable lesson. When two credible sources disagree, the disagreement is the datum. An aggregator must pick one number and therefore transmits a false precision. A research platform that records the range and the reason transmits a true uncertainty — and in this case the gap between the two methodologies pointed directly at the finding that matters, which is that dollars and counts have decoupled.


Case 2 — Humanoid robots: how a prospectus and an S-4 falsified a narrative that video had established

What was claimed. Humanoid robots are entering productive industrial deployment. "Deployment trackers" returned by search report thousands of units in the field. The category carried roughly $100bn+ of aggregate valuation: Figure AI at $39bn (September 2025), Unitree peaking near $66bn, XPeng's Dogotix raising $900m at $6.3bn. The evidence base circulating publicly was video demonstrations and partner logos.

What the evidence showed. Three disclosure events in 2026 made the category checkable for the first time — and every one of the damaging numbers comes from a party with an incentive to talk the category up.

Disclosure Date What it revealed
Unitree STAR Market prospectus / IPO 2026-08-19 5,632 humanoids shipped cumulatively 2023–2025, over 5,500 of them in 2025. Humanoid revenue RMB 868m, 51.78% of total. Less than 10% of 2025 revenue from industrial applications. The G1 sells at $13,500 — a research, education and entertainment price point
Agility Robotics Form S-4 (Churchill Capital Corp XI) 2026-09-04 $1.8m of 2025 revenue against a $140m operating loss. The most-cited humanoid logistics deployment in the world generates less revenue than a single mid-sized systems integrator
Tesla statements January 2026 Zero Optimus units performing useful work, against a promise of 10,000 in 2025. Musk declined to give a 2026 target, calling output "quite slow" and "impossible to predict"

Corroborated by the industry's own trade body: IFR states that mass adoption as universal factory helpers or in households "will not happen within the near- and medium-term future." And by a pre-IPO sell-side note: HSBC warned that "the surge in shipments for robot makers could be illusionary."

The market then confirmed it. T-13-16: Unitree fell from an RMB 1,100 first-day high to RMB 513.93 on 2026-09-09 — 53% off the high, 39% off the first-day close, erasing roughly $35bnwhile nothing changed operationally.

How the methodology caught it. The decisive act was definitional, not investigative. Search results for humanoid deployment counts are dominated by trackers that count announced pilots, letters of intent, demonstration units and units produced as "deployments." Filings count revenue. The research contract's source-tier rules exclude undated, unauthored content-farm pages on sight, which removed the entire tracker layer from the evidence base and left only filings — at which point the answer was arithmetic.

The second act was claim typing. "Every industrial company will become a robotics company" and "GR00T N2 succeeds at new tasks more than twice as often" are recorded as marketing (T-13-20) and may never be rendered without the label. The Unitree and Agility figures are fact. Placing them in the same record makes the gap unavoidable.

Which scoring dimensions moved. T-13-19 is the most instructive score vector in the database:

velocity 4   capital 5   strategic_importance 4    ← attention and money are real
adoption 1   revenue 1   customer_demand 1         ← nobody is buying
persistence 2   technical_maturity 2
evidence_quality 5   source_diversity 5            ← E = 1.00, cap = 100
composite 51.1 (uncapped)

The cap did not bind, and that is the point. The evidence is excellent — because prospectuses exist. The score is low because adoption and revenue are 1 and 1. A model that let attention and capital carry the composite would have scored this trend in the 80s. On T-13-11 (the listings themselves) regulatory_impact rises to 4, because Chinese regulators informally raised the bar for humanoid IPO candidates, requiring recurring revenue or progress toward profitability — a second-order consequence of the same disclosure.

Generalisable lesson. An IPO or an S-4 is a trend-research event, not just a capital-markets event. Public listing forces disclosure, and disclosure is what makes a narrative falsifiable. The practical rule this case generates: treat every demonstration video as marketing until it is accompanied by an independently verified paying deployment — and watch component orders, not demos. Harmonic Drive's consolidated bookings and Schaeffler's 2027 strain-wave line (T-13-10) are better leading indicators of genuine humanoid volume than any humanoid company's announcements.

Stated honestly: unit figures for Figure, Apptronik, 1X and AGIBOT are entirely undisclosed. The installed base is unknown, not known to be small.


Case 3 — The EU AI Act high-risk deadline: a compliance market sold against a date that slipped sixteen months

What was claimed. The EU AI Act's high-risk obligations for employment-context AI — recruitment, performance evaluation, worker monitoring, promotion, termination — would bind on 2026-08-02. This was the single most-cited compliance deadline in vendor marketing across at least five sectors, and a large volume of AI-governance tooling was sold against it.

What the evidence showed. The Digital Omnibus on AI entered into force 2026-07-27, six days before the deadline, deferring high-risk employment obligations to 2027-12-02 — a ~16-month postponement pending harmonised standards. The AI Act's transparency obligations did take effect on 2026-08-02 as scheduled, and the two are routinely conflated.

How the methodology caught it. The definitions framework requires that regulation be tracked through four distinct stages — proposed, enacted, in force, enforced — and states that vendor marketing routinely collapses them. Applying that discipline produced five separate corrections across five sectors, all from one legal event:

Sector Record What the deferral invalidated
01 T-01-16 (51.5) The compliance wave itself. "Any vendor still selling against the August 2026 date is selling a stale deadline"
02 T-02-20 The regulatory premise under agentic-AI governance marketing
04 T-04-19 The one regulatory hook that might have touched analyst displacement
09 T-09-19 "Removing the compliance deadline that much 2025-26 vendor marketing was built on"
14 / 23 T-14-19, T-23-20 Statutory constraints on AI in production and design are lighter than assumed

The methodology then generalised it rather than treating it as a one-off. Two independent precedents were retrieved and recorded:

  • T-16-18: third-party cookie deprecation moved 2022 → 2023 → 2024 → 2025 and was ultimately abandoned as a forced migration. Six years of product roadmaps on a date that never bound. The dossier calls it "the sector's best-documented false positive and the most useful calibration case a trend platform can carry."
  • T-19-09: FSMA Section 204 food traceability moved from 2026-01-20 to a statutory direction not to enforce before 2028-07-20 — a 30-month slip.

Which scoring dimensions moved. T-01-16 carries regulatory_impact 5 — the event is entirely regulatory — against capital 1, customer_demand 1 and adoption 2. persistence is 3, not higher, because the deferral could itself be reversed. source_diversity is 2 (official EU instruments are authoritative but few), which sets E = 0.60 and a cap of 76; the composite of 51.5 sits below the cap, so the cap is not doing the work — the low adoption and capital scores are.

Generalisable lesson. A deadline is a forecast, and forecasts about legislatures have a measurable base rate. In this research cycle alone: the EU AI Act slipped 16 months, FSMA 204 slipped 30 months, cookie deprecation slipped four times and was abandoned, and the US FOP "Nutrition Info box" rule has produced no final rule fourteen months after comments closed. Any revenue forecast, budget or trend thesis anchored on a future compliance date should be treated as low-persistence by construction and discounted against that base rate. Contrast §Case 4: where a court acted, the change landed inside a quarter.


Case 4 — The tariff-refund windfall: one court ruling, three sectors' margins, and a 2027 cliff

What was claimed. US retailers, brands and industrials reported materially expanded gross margins in 2026. Coverage read this as operating improvement — pricing power, sourcing discipline, cost control.

What the evidence showed. A large share of it is a non-recurring legal windfall. The Supreme Court voided the IEEPA reciprocal and trafficking tariffs on 2026-02-20 (6–3) in Learning Resources v. Trump and Trump v. V.O.S. Selections, holding that IEEPA does not authorise tariffs of indefinite scope. Penn Wharton estimated up to $175bn of refundable duty; CBP had certified roughly $107bn of refunds by 2026-08-21 (T-09-01).

The refunds then appeared in cost of sales across three sectors:

Company Refund booked Margin effect Source record
Target $994m 3.7pp of gross and operating margin; ~$1.65 of FY EPS T-12-03
Walmart not separately sized 96bp gross-margin gain, attributed "primarily" to refunds T-12-03
NIKE $986m ~900bp of fiscal Q4 gross margin; $0.52 of $0.72 EPS T-23-01
lululemon $134.5m + $4.1m interest 560bp of gross margin; $0.86 of EPS T-23-01
e.l.f. Beauty not separately sized ~1,050bp of a 1,400bp gross-margin expansion T-23-01
PUMA €15.4m booked, €33.8m applied for T-23-01
Estée Lauder $38m T-23-01
adidas excluded from guidance potential $250–300m T-23-01
Starbucks "substantially all" requested refunds received "largely offset" tariffs in the first three fiscal quarters T-19-18

The decisive figure is lululemon's. T-23-16: reported gross margin rose 200bp to 60.5% — but the refund plus interest was worth 560bp, so the underlying margin fell roughly 360bp, in a quarter when comparable sales fell 9% and Americas comps fell 12%. A company whose margin rose reported an underlying margin that contracted.

How the methodology caught it. Three steps an aggregator cannot take:

  1. Connecting a February court ruling to an August earnings line. The two events are seven months and two sectors apart, and neither press release makes the connection explicit. The dossier for sector 12 names this directly as "most overlooked... because the refund arrived as good news inside good quarters."
  2. Cross-sector propagation. The same instrument was traced into retail (T-12-03), fashion and beauty (T-23-01), industrials (T-09-01) and food service (T-19-18). A sector-siloed process finds four unrelated margin stories.
  3. The persistence dimension forces the 2027 question. Scoring persistence at 2 requires the analyst to state what happens when the effect stops.

Which scoring dimensions moved. T-12-03: revenue 5, regulatory_impact 5, strategic_importance 5, velocity 5, E = 1.00 — and persistence 2. T-23-01 is more extreme still: revenue 5 and adoption 5 against capital 0 and customer_demand 0. A record with maximum revenue impact and zero customer demand is the signature of a windfall, and the score vector says so on its face.

Generalisable lesson. Ask every 2026 margin expansion what the ex-refund number is. More generally: when a legal or policy event transfers cash into cost of sales, report the adjusted series and put a date on the reversal. The reversal mechanism here is the simplest in the entire research programme — arithmetic. Every company above faces a 2027 gross-margin comparison it cannot repeat, and in fashion it lands simultaneously with a forecast 22% rise in the US upland cotton farm price against fifteen-year-low world ending stocks (T-23-14).


Case 5 — The autonomous AI SOC: $300m+ of funding and zero independent efficacy evidence

What was claimed. Agentic security-operations products that triage, investigate and close alerts without a human, marketed through 2026 by Torq, 7AI, Qevlar AI, Conifers, Dropzone, Radiant and every platform incumbent. The capital is entirely real: Torq $140m at a $1.2bn valuation (January 2026), 7AI $130m (December 2025), Qevlar $30m (March 2026) — roughly $300m+ disclosed in twelve months.

What the evidence showed. A deliberate search for adoption or efficacy evidence in September 2026 returned exclusively vendor blogs, vendor-sponsored buyer's guides and consultancy content marketing. No neutral study. No regulator dataset. No peer-reviewed evaluation of autonomous triage accuracy in production was locatable.

Meanwhile, the one independently measured defensive metric that moved in 2026 moved the wrong way: median patching time rose from 32 to 43 days and KEV remediation fell from 38% to 26% (Verizon DBIR 2026).

Three further disconfirming observations were recorded:

  • Vendor alert-reduction percentages measure workload inside the tool, not detection efficacy, and are trivially gamed by suppressing alerts.
  • The category's direct historical analogue, SOAR, made the identical promise in 2017 and failed on the identical problem — and is rarely mentioned in current marketing.
  • Per OWASP, prompt injection drives most agentic AI failures in production. A defensive agent that reads attacker-controlled alert content is itself a prompt-injection target.

How the methodology caught it. This is the clearest case of absence recorded as a finding rather than as a gap in research effort. The record states it explicitly: "a search for neutral adoption evidence returns only vendor content — an absence that is itself the finding, and the reason evidence_quality is scored 1."

Two contract rules did the work. First, marketing claims may not be rendered unlabelled, which stripped the category's entire evidence base. Second, the evidence cap made the result arithmetic rather than editorial.

Which scoring dimensions moved.

evidence_quality 1   source_diversity 1   → E = 0.20 → cap = 40 + 60(0.20) = 52
capital 4   velocity 3   social_impact 4
adoption 1   revenue 2   persistence 2   regulatory_impact 1
composite 50.4   confidence low   verification_status unverified

E = 0.20 is the lowest evidence factor in the database, and it produces the lowest cap: 52. However loud this category becomes, and however much capital it raises, it cannot score above 52 until someone publishes a false-negative rate. That is enforced arithmetic, not editorial intention — and it is the cleanest demonstration of the cap's purpose in the entire programme.

Generalisable lesson. Search for the disconfirming evidence specifically, and record its absence with the same rigour as a positive finding. The operational form for a buyer: a CISO who cannot obtain a false-negative rate and an escalation-accuracy figure is buying a labour substitution with no measured error rate, in a function where the error is a missed breach. And the structural warning: displacing the entry tier removes the training pipeline for the senior analysts these systems still require.


Case 6 — Data-centre load forecasts: revised 69% up in a year, and already deflating in three jurisdictions

What was claimed. Unprecedented electricity demand growth driven by AI. Every major North American system operator revised its decade-ahead forecast sharply upward between late 2025 and mid-2026.

What the evidence showed — both halves.

The upward revision is real and enormous. T-05-01 (composite 90.2, E = 1.00): NERC's 2026 Long-Term Reliability Assessment raised the ten-year summer peak increase to 224 GW, 69% above the 132 GW projected a year earlier, and called it the highest compound growth rate since NERC began tracking in 1995. PJM now forecasts 3.6% annual summer peak growth against 0.3% in its 2021 forecast — a twelve-fold change in the planning assumption. ERCOT's preliminary 2026–2032 forecast reaches 367,790 MW by 2032 against an all-time actual peak of 85,508 MW.

And the first credible down-revisions have already arrived. T-05-18 (62.9, E = 1.00):

  • PJM cut its summer 2027 forecast by ~4 GW and 2028 by 4.4 GW (2.6%) after applying stricter vetting to planned data centres.
  • AEP Ohio halved its data-centre load forecast.
  • Exelon stated that only 22% of its 65 GW pipeline through 2040 is likely to materialise.
  • ERCOT suspended its Batch Zero large-load study process on 2026-08-07 against roughly 474 GW of pending requests, about 90% data centres — more than five times its historical peak demand — and ERCOT itself said it expects its forecast to be higher than actual load growth.
  • CenterPoint's Houston-area requests rose from 1 GW to 25 GW within twelve months, which is what made the duplication problem visible.

How the methodology caught it. Four discriminations, none of which an aggregator makes:

  1. Both records coexist. The database holds the 90.2-scoring record that load forecasts were re-rated upward and the 62.9-scoring record that they are deflating, because both are true and neither is the other's refutation.
  2. Scope discipline. T-01-20 separates the national claim from the regional fact: EIA forecasts 1% US demand growth in 2026 and 3% in 2027 (baseline 1.9%/2.5%), while ERCOT forecasts ~10%/yr and PJM 3%. "Framing a procedural and regional bottleneck as a national generation shortfall misdirects capital toward generation and policy toward emergency supply measures, when the binding constraint is queue position, cost allocation and local permission."
  3. Request ≠ capacity. Queue volume is optionality purchased, against a ~13% historical completion rate (T-05-05). Applying Exelon's 22% to JLL's 66 GW pipeline (T-18-20) describes an entirely different real-estate market from the headline.
  4. Who is doing the correcting. The record notes that PJM, ERCOT and Exelon — the operators now vetting — are the same ones whose forecasts justified the spending. The correction is coming from inside. The instruction that follows is precise: watch the vetting criteria, not the headline.

Which scoring dimensions moved. T-05-01 scores velocity 5, capital 5, revenue 5, breadth 5, depth 5, customer_demand 5, regulatory_impact 5 and E = 1.00 — and persistence 3. The persistence score is the finding: the forecast itself is unstable. The record's own why_it_matters says so — "decision-makers are committing thirty-year assets against a number that moved twelve-fold in five years and was revised 69% in a single year. The forecast is now the most consequential and least reliable input in the industry."

T-05-18 scores evidence_quality 5 and source_diversity 5 (E = 1.00) against customer_demand 1 and persistence 2, and carries an explicit honesty caveat: PJM simultaneously raised its long-term growth rate from 3.1% to 3.6%. This is deflation at the margin, not reversal, and the record refuses to overstate.

Generalisable lesson. A forecast that has been revised by 69% in one year is not a planning input; it is a variable to be hedged. The portable discriminations are: separate national from regional; separate requested from contracted; and note that the vetting criteria are not published, so the correction is not auditable — which is itself a finding to carry forward.


Case 7 — The DPI drought: record paper returns, almost no cash

What was claimed. Venture capital recovered in 2026. Record H1 deployment, record exit value, a 17.1% one-year horizon IRR among the highest of any private-capital strategy, and fund asset values up 21.6% on AI markups. Q2 2026 global exit value reached $1.9tn, exceeding the previous annual record.

What the evidence showed. The cash tells the opposite story.

  • 2021-vintage VC funds average 0.05x DPI — the lowest five-year DPI multiple this century (T-11-01).
  • Net cash flow to LPs has been negative $202bn since 2022. Since the start of 2022 US VC managers have called 1.6x more capital than they have distributed, against 2012–21 when distributions exceeded calls by 1.3x.
  • 2019 and 2020 vintage median DPIs are barely above zero; fewer than half of those funds have returned any capital; under 20% of 2017–18 funds have reached 1x DPI.
  • The exit value is four companies: 93.5% of 2026 venture exit value (T-11-19). 44 US VC-backed IPOs priced by late July against 950+ unicorns — under 5% of the eligible population — with SpaceX −30.3% and Cerebras −34.7% from debut against a Renaissance IPO Index up 12.7%.
  • The consequences compound downward: seed→Series A graduation fell from 55%+ to 16% for the 2024 cohort (T-11-03); first-time fund formation is on pace for its lowest year since 2016 with median time between closes at 1.7 years (T-11-17); and the liquidity valve itself is being tested — Blue Owl Capital Corp. II closed quarterly redemptions on 2026-02-18, the $33bn Cliffwater fund received requests for 14% against a 7% cap, and non-listed BDC redemptions hit 4.71% of NAV, nearly tripling quarter on quarter (T-11-16).

How the methodology caught it. Three mechanisms:

  1. IRR is a marked number; DPI is a cash number. The methodology treats an IRR computed on unrealised, self-reported marks during a period of AI-driven markups as "a restatement of the markup, not independent evidence about it."
  2. Counts alongside sums, everywhere. Applying the rule produced the graduation collapse, the deal-count collapse and the fund-formation collapse as separate, mutually reinforcing records.
  3. Recording an absence as a finding. The dossier records that no independent, regulator-published, fund-level performance dataset exists — every DPI figure comes from a commercial vendor with a different self-selected sample — and notes why nobody says so: "the vendors are excellent at marketing their data and nobody's business model depends on pointing out that it is not audited."

Which scoring dimensions moved. T-11-01's vector is diagnostic:

adoption 5   capital 5   breadth 5   depth 5   geographic_spread 5
persistence 5   strategic_importance 5   evidence_quality 5   source_diversity 5
revenue 2                                   ← the entire finding
composite 84.8, E = 1.00, uncapped

A record scoring 5 on nine dimensions and 2 on revenue is precisely the shape of "large, real, well-evidenced, and not producing cash." Four separate data organisations using different samples agree on the direction, which is what carries source_diversity to 5 despite the absence of an audited benchmark.

Generalisable lesson. Distinguish the metric the industry reports from the metric its customers can spend. IRR, TVPI, NAV, TVL, GMV, bookings and "represented asset value" are all marked or gross measures; DPI, cash distributions, audited revenue and posted margin are not. The same discrimination appears in four other sectors in this research — T-15-17 (Roblox bookings vs recognised revenue), T-21-17 (DeFi TVL vs Coinbase's audited blockchain rewards revenue, −42%), T-14-20 (box-office revenue vs admissions) and T-25-04 (travel revenue vs RPK). It is the single most transferable analytical move in the programme.


Case 8 — The creator-economy market size: a number with no traceable methodology

What was claimed. The creator economy is worth a specific multi-hundred-billion-dollar sum. The figure appears in strategy decks, investment memos and policy submissions.

What the evidence showed. In this research cycle, no such figure could be traced to a nameable methodology or an identifiable modeller. What can be sourced is an order of magnitude smaller and consists of platform-disclosed payouts:

  • YouTube: $100bn over four years — and the disclosure covers creators, artists and media companies, not independent creators alone. That is ~$25bn/year, gross, across a population much wider than "creators."
  • Roblox: $1.5bn in 2025, against $923m in 2024 — +62%, and a genuine, auditable year-over-year growth series.

The record also captured the laundering mechanism itself, which is the more valuable finding: searching for YouTube's payout figure in September 2026 returns, among the top results, a February 2024 article stating "$70 billion over three years" with no recency signalling. And CNBC reported the same $100bn over four years in September 2025, sixteen months before the January 2026 CEO letter repeated it — so the figure is rounded and no annual growth series can be derived from it.

How the methodology caught it. Four rules, applied in sequence:

  1. Name the modeller or the figure is unusable. An estimate requires a named modeller. No modeller could be named, so the figure could not be entered at any claim type.
  2. Tier-C content-farm sources are excluded on sight. The search landscape for this statistic consists almost entirely of pages meeting that definition.
  3. Check the publication date of every result before using it (macro addendum 2, statistical trap 3). This is what surfaced the date-laundering mechanism.
  4. Rounded cumulative disclosures cannot be differenced (trap 4). This is what prevented the platform from manufacturing a growth rate out of $70bn → $100bn.

Which scoring dimensions moved. T-17-19 is the lowest-scoring record in the entire 500-record database at 28.7, and its vector is unlike any other:

velocity 1   capital 0   adoption 0   revenue 0   technical_maturity 0
evidence_quality 1   source_diversity 2   → E = 0.30 → cap = 58
persistence 4   geographic_spread 4   breadth 3
verification_status: unverified

persistence 4 is the deliberate anomaly. The trend being scored is not a market — it is a statistic that will not die. It is durable, geographically widespread and has no capital, no adoption, no revenue and no technical content, because it is a number rather than a phenomenon. The record states this plainly: "this record concerns a statistic, not a product."

Generalisable lesson. Refuse any market sizing that does not name a modeller, a date and a methodology — and substitute the auditable floor. Doing so here changed the answer by roughly an order of magnitude, which changes every conclusion built on it: creators making career decisions on inflated opportunity estimates, investors underwriting to unsourceable TAMs, and policymakers reasoning from bad numbers.

Recorded honestly: sector 17 completed with a constrained search budget and the record carries the explicit caveat that "absence of evidence within one constrained research cycle is not proof that a rigorous estimate does not exist... This record should be revisited at full search budget." The finding is "we could not trace it," not "it does not exist."


9. What the eight cases have in common

# Case The discriminating act The general form
1 Crunchbase vs KPMG Recorded the contradiction rather than resolving it Disagreement is data
2 Humanoids Excluded trackers; used filings Definition beats volume
3 EU AI Act Tracked four regulatory stages separately A deadline is a forecast
4 Tariff refunds Connected a February ruling to an August earnings line across three sectors Adjust for the one-off, and date the reversal
5 Autonomous SOC Recorded an absence as a finding Search for disconfirmation specifically
6 Load forecasts Separated national from regional, requested from contracted Scope before magnitude
7 DPI drought Separated marked returns from cash returns Which metric can the customer spend?
8 Creator economy Demanded a named modeller; checked result dates No modeller, no number

Two properties run through all eight.

First: in every case the disconfirming evidence was free and public. CPUC filings, the Agility S-4, the Unitree prospectus, NERC's LTRA, PJM's forecast filings, Target's and NIKE's earnings releases, the Verizon DBIR, KPMG's and Crunchbase's own published methodologies, YouTube's own letter. Nothing required privileged access. The gap these cases close is an attention-allocation gap, not an information gap.

Second: in every case the error was relational, not factual. No source lied. The aggregator-defeating step was always the same kind of move — comparing two numbers that are normally read separately (uploads and streams; bookings and revenue; announcements and put-in-place; IRR and DPI; ceiling and obligation; national and regional; reported margin and refund). That is the methodology's actual product: not better facts, but the relation between facts that nobody is paid to compute.

Research provenance
Source artifact
04-analysis/44-case-studies.md
Corpus date
15 September 2026
Prepared for this site
16 September 2026
Site publication
18 September 2026
Verification
Inherited; not fully rechecked