SASIGNAL ATLASCross-industry intelligence / Research desk
SIGNAL ATLAS / RESEARCH DESK

3 — Cross-Industry Taxonomy

Method · Cross-industry · Original Phase 1 research

3 — Cross-Industry Taxonomy

Phase 1 deliverable · research date 2026-09-15


3.1 What the taxonomy has to carry

The taxonomy has three jobs, and they pull against each other.

  1. Comparability. A semiconductor trend and a fashion trend must be classified by the same rules, or the cross-sector ranking in §2 is arithmetic on incommensurable things.
  2. Enum stability. Every facet below becomes a database column with a constrained domain. Once 500 records — and eventually millions — are stamped with a value, that value cannot be silently redefined without invalidating every historical query.
  3. Extensibility. Phase 2 requires user-defined taxonomies. Those must sit on top of the canonical vocabulary without contaminating it.

The governing design rule is: a facet earns its place only if a different value would change a decision. Facets that merely describe are dropped; facets that discriminate are kept and controlled.

This document is not aspirational. It describes the taxonomy that the 500 scored trend records, 994 entity records and 603 registered sources of the Phase 1 seed database actually instantiate — and it marks explicitly, in §3.5, the nine places where the records reveal that the specification and the practice have diverged.


3.2 Trend types: seven families, not a flat list of thirty

The brief supplies thirty trend types. Left flat, thirty categories is a menu, not a taxonomy: it offers no guidance on which categories are alternatives to each other, and it makes a technology trend and a regulation trend look like the same kind of object when they require completely different evidence.

The grouping principle: two trends belong to the same family when the same kind of evidence would falsify them. This is the only grouping that does analytical work. Alternatives were considered and rejected — grouping by subject matter reproduces the industry axis, and grouping by "PESTLE"-style macro categories produces buckets that no detection signal maps onto.

Family Canonical types Question it answers What falsifies a member Primary evidence class Lead signal
A · Capability technology, products, design, infrastructure What has become technically possible or economically cheap? A demonstrated capability that does not convert into deployment Patents, standards, benchmarks, shipped SKUs, technical filings Patents, open-source activity, academic research
B · Commercial model business_models, pricing, platforms, distribution, marketing How is value captured, and by whom? Disclosed take-rate, ARPU, margin or channel mix that does not move Earnings, price lists, contract terms, platform policy pages Price changes, product launches
C · Capital startups, venture_capital, funding, investment, m_and_a Who is paying, on what terms, and did anyone get paid back? Capital deployed without realisation Regulatory filings, funding databases, exit and distribution data Funding growth, M&A, new business formation
D · Rules and power regulation, policy, geopolitics, privacy, security What does the state permit, require or forbid? Stage slippage: proposed ≠ enacted ≠ in force ≠ enforced Primary legal instruments, dockets, enforcement actions, procurement Regulatory attention, government procurement
E · Physical economy supply_chains, manufacturing, climate, sustainability What do the atoms, the capacity and the constraints do? Announcement volume exceeding docketed or physically observable throughput Customs data, capacity registries, lead times, installation counts Supply-chain shifts, price changes
F · People and demand consumer_behaviour, culture, labour_and_workforce, education, health, media What do people actually do, and who does the work? Stated intent diverging from behavioural data Transaction, payroll, enrolment, traffic and usage data Hiring, customer adoption, consumer behaviour
G · Risk risk What is the failure mode, and who is holding it? Cross-cutting — a risk trend is falsified by the absence of the transmission channel it posits All of the above, inverted — (derived)

Thirty types, seven families, each type placed exactly once.

Why the families behave differently in practice, with records from the database:

  • A · Capability trends fail at the deployment boundary, not the research boundary. T-13-19 (Humanoid robots doing useful paid work at scale) has evidence_quality 5 and source_diversity 5 — the evidence is excellent — and still scores 51.1 composite, because the evidence documents 5,632 cumulative Unitree shipments and $1.8m of Agility Robotics revenue against a $140m operating loss. Capability is advancing; deployment is not. Both statements are true simultaneously, which is the defining property of this family.
  • B · Commercial-model trends are the most numerous and the least self-evident. Normalised, business_models carries 27.7% of all type assignments. T-24-03 (Music revenue growth is coming from price, not from listeners) is the shape: the headline metric rose, and the composition of the rise is the finding.
  • C · Capital trends are falsified by realisations, not by deployment. T-11-01 (The DPI drought: rising marks, almost no cash returned) — 2021-vintage DPI at 0.05x, net LP cash flow since 2022 at −$202B — is the reference record for the entire family. Every capital trend must be read against it.
  • D · Rules-and-power trends must carry their stage. The EU AI Act's high-risk obligations for employment AI were deferred from 2026-08-02 to 2027-12-02. Any record in this family that does not distinguish proposed / enacted / in force / enforced is unusable, because vendor marketing collapses those four states by default.
  • E · Physical-economy trends are the family where announcements and throughput diverge most reliably. Sector 05's hydrogen precedent — "announcement volume far exceeding docketed projects" — is the canonical failure. T-20-01 and the CDR record of 1.68 Mt delivered against 49.47 Mt contracted is the 2026 instance.
  • F · People-and-demand trends diverge between stated intent and behaviour, always in the same direction. T-25-20 records travel-intent surveys pointing the opposite way from every primary volume series (global RPK +0.2%, US passengers −1.4%, Amadeus bookings −3.7%, Japan arrivals −6.8%).
  • G · Risk is a modifier as much as a family. Nine assignments in 500 records use it as a primary type; it does more work through the risks[], losers[] and counter_trends[] arrays than through the type facet.

What the 500 records actually used

tags.trend_type was specified in the research contract as an open list. Analysts treated it as free text, and produced 75 distinct values across 1,151 assignments. Normalising those 75 against the canonical vocabulary collapses them to 34 types with no residue.

Normalised type Assignments Share Status
business_models 319 27.7% canonical
regulation 208 18.1% canonical
technology 190 16.5% canonical
market_structure 50 4.3% needed, not in the brief
supply_chains 42 3.6% canonical
infrastructure 37 3.2% canonical
consumer_behaviour 33 2.9% canonical
capital_markets 32 2.8% needed, not in the brief
macroeconomics 29 2.5% needed, not in the brief
pricing 27 2.3% canonical
labour_and_workforce 24 2.1% canonical
policy 22 1.9% canonical
geopolitics 17 1.5% canonical
narrative 14 1.2% needed, not in the brief
standards 12 1.0% needed, not in the brief
security 11 1.0% canonical
risk 9 0.8% canonical
measurement 9 0.8% needed, not in the brief
commodity_cycles 7 0.6% needed, not in the brief
culture, manufacturing, distribution 6 each 0.5% canonical
m_and_a, investment, marketing, trade 5 each 0.4% 3 canonical, trade new
sustainability 4 0.3% canonical
funding, media, platforms, products, demographics 3 each 0.3% 4 canonical, demographics new
climate, health 1 each 0.1% canonical

Three findings, all auditable against the file:

  1. Three types carry 62.3% of all assignments. business_models, regulation and technology are doing the work of a general-purpose bucket. A facet whose top three values cover nearly two-thirds of the corpus has weak discriminating power and must be used with the family layer, not instead of it.
  2. Five of the brief's thirty types have zero instances: design, education, privacy, startups, venture_capital. This is not because the subject matter is absent — sector 22 is education, sector 11 is venture capital, and age-assurance and data-protection trends are numerous. It is because those five are industry labels wearing a trend-type costume. An analyst classifying a venture-capital trend correctly reached for business_models or capital_markets, because "venture capital" describes whose trend it is, not what kind of change it is. Recommendation: retire startups, venture_capital and education as trend types — they are already captured by industry_id — and retain design and privacy, which are real change-kinds that this 25-sector roster simply under-samples.
  3. Nine types were needed and are not in the brief, covering 14.0% of assignments: market_structure, capital_markets, macroeconomics, narrative, standards, measurement, commodity_cycles, trade, demographics. Two deserve comment. narrative (14 assignments) is a genuine innovation: it marks trends whose object is a claim in circulation rather than a change in the world — T-24-19 ("AI music is taking over streaming"), T-25-19 (Agentic AI disintermediating online travel agencies), T-16-19 ('AI search is destroying demand capture' as a budget argument). All three are classification: overhyped. A platform whose mission is calibration needs a type for trends that exist mainly as assertions. measurement (9 assignments) marks trends about the data itself — T-11-01's DPI problem, T-17-04's displacement of market-size estimates by platform-disclosed payouts. Both should be promoted to canonical.

Canonical trend-type vocabulary, v1.0: 32 values — the brief's thirty, minus startups, venture_capital and education, plus market_structure, capital_markets, macroeconomics, narrative, standards, measurement — organised into the seven families above, with commodity_cycles, trade and demographics held in the proposal queue (§3.6.1) pending a second corpus.


3.3 The facet system

Fourteen facets. Every one is a column; every one has a defined domain, cardinality and owner.

# Facet Card. Domain size Assigned by Present in seed DB
1 industry_id single 25 automated (routing) + editorial confirm yes
2 subindustry single uncontrolled (404 strings) editorial yes, as free text
3 geography[] multi 3-tier, ~270 automated (NER) + editorial yes, uncontrolled
4 customer_type[] multi 4 editorial yes
5 companies[] multi entity FK automated (NER + resolution) yes, as free text
6 products[] multi product FK automated + editorial yes, as free text
7 technologies[] multi ~200 automated + editorial absent
8 market_stage single 8 editorial yes, duplicated
9 time_horizon single 4 editorial yes
10 confidence single 3 computed, editorial override yes
11 direction single 4 editorial yes
12 impact derived 5 bands + 15 dimensions computed partially
13 source_type single (per source) 27 editorial at registration yes
14 evidence_quality single 0–5 + status enum editorial, arithmetically capped yes

3.3.1 industry_id — single-valued, automated with editorial confirmation

Allowed values: the 25 two-digit identifiers from 02-industry-roster.md, 0125. IDs are stable identifiers, not ranks; they never change meaning. A 26th sector (telecom and connectivity, the roster's flagged first addition) gets 26, never a reassignment.

Cardinality: exactly one. This is the taxonomy's most consequential constraint and §3.5.1 argues it is wrong.

Assignment: a classifier routes on entity mentions and source registry membership; an editor confirms. In the seed database, assignment was editorial by construction — each of the 25 sector agents wrote exactly 20 records, producing a perfectly uniform 20-per-sector distribution.

Worked example: T-18-01 Data-centre real estate as the only structurally tight major property typeindustry_id: "18" (real-estate-construction). Its subject is a data centre; its market is commercial property, and the roster assigns data-centre real estate to 18 while data-centre compute sits in 01 and interconnection in 05. The routing rule is market of the buyer, not subject of the sentence.

3.3.2 subindustry — single-valued, editorial

Allowed values: the 6–12 subindustries enumerated in each dossier's §2, giving a canonical set of roughly 200 across 25 sectors.

Reality in the seed database: 404 distinct strings for 500 records. Only 63 are used more than once; 341 are used exactly once. 107 are compound values"Neoclouds / GPU specialists", "Horizontal SaaS / pricing and packaging", "Recorded music / streaming platform operations". The facet is currently a free-text note field.

Fix: promote the dossier §2 lists to the controlled vocabulary, make the facet multi-valued (the compound strings are the corpus telling us it needs to be), and preserve the original string in subindustry_raw for migration audit.

Worked example: T-01-01industry_id: "01", subindustry: "Hyperscale AI cloud". Clean. T-02-19subindustry: "Horizontal SaaS". Also clean. The failure mode is T-23-19"Apparel sizing; beauty adjacency", which is two subindustries in one string.

3.3.3 geography[] — multi-valued, automated with editorial confirmation

Full vocabulary in §3.4. Cardinality in the seed database runs 1–9 tokens per record (median 3); 114 records carry a single token and 16 carry seven or more.

Worked example: T-09-04 China's direct share of US imports collapses while connector economies absorb the flow["US","CN","VN","MX","IN","TW","KR","TH"]. Eight ISO-3166-1 alpha-2 country codes, no blocs, no regions. This is the facet working correctly: the trend's mechanism is a specific set of bilateral flows and the geography facet names them.

3.3.4 customer_type[] — multi-valued, editorial

Allowed values: b2b, b2c, b2g, b2b2c. Four values, closed set, no drift observed across 500 records — the only facet in the seed database with perfect vocabulary discipline.

Distribution: b2b 408, b2c 200, b2g 147, b2b2c 41. Cardinality: 229 records single-valued, 246 two-valued, 25 three-valued.

Assignment: editorial. Automation is unreliable here because the distinction that matters — b2b2c versus b2c — depends on who holds the customer relationship, which is a business-model judgement rather than a textual feature.

Worked example: T-15-04 UGC creator payouts are a low-twenties percentage of bookings["b2b2c","b2c"]. Roblox sells to players (b2c) and the trend concerns the economics of the creators who sit between platform and player (b2b2c). Tagging it b2c alone would lose the entire subject.

Note on b2g: 147 of 500 records — 29% — carry it. That is far higher than a consumer-technology framing would predict, and it is the single strongest argument in the data for treating government procurement as a first-class detection signal (§7).

3.3.5 companies[] — multi-valued, automated extraction with entity resolution

Allowed values: foreign keys into the entity table, E-<NN>-<NN>.

Reality in the seed database: free-text strings. 1,689 distinct company strings across 3,330 mentions, of which only 272 — 16.1% — match an entity record's name exactly. "Meta Platforms" (20 mentions) and "Meta" (11 mentions) are two strings for one company. The entity table itself is not deduplicated: 994 records resolve to 924 distinct names, with OpenAI appearing as nine separate entity records (one per sector that needed it) and Anthropic as five.

Fix, in order: (a) a global entity registry keyed on legal entity with LEI/CIK/ticker where available; (b) companies[] becomes an array of entity IDs; (c) the display string is resolved at read time, so a company rename propagates without a data migration.

Worked example (target state): T-03-01 Memory supercyclecompanies: ["E-GLOBAL-MICRON","E-GLOBAL-SKHYNIX","E-GLOBAL-SAMSUNG","E-GLOBAL-KIOXIA"], resolving to Micron Technology, SK hynix, Samsung Electronics, Kioxia — each carrying its own ticker, HQ country and sector memberships, so that "which trends touch Samsung" becomes a query rather than a string search.

3.3.6 products[] — multi-valued, automated with editorial confirmation

Allowed values: foreign keys into a product table, each product owned by exactly one entity.

Reality: 1,669 distinct strings across 1,851 mentions — a 90% singleton rate, meaning the facet is currently prose. Cardinality 1–6, median 3.

Why it still matters: product-level granularity is what separates a real capability trend from a category narrative. T-03-01 names HBM3E, HBM4, DDR5; the trend is falsifiable against those three part families and their contract pricing. A record that named only "memory" would not be.

Worked example: T-06-01 Incretin franchise trades price for volume["Mounjaro / Zepbound (tirzepatide)", "Ozempic / Wegovy (semaglutide)", "oral Wegovy"]. Note that even here the strings bundle brand pairs and molecules. The target schema separates product (brand), molecule_or_platform, and owner_entity_id.

3.3.7 technologies[] — multi-valued, automated with editorial confirmation

This facet does not exist in the seed database. No trend record has a technology field and no tags.technology key appears in any of the 500 records. Technology is currently encoded twice, badly: once as trend_type: "technology" (190 normalised assignments, which says only that it is a technology trend, not which technology), and once as unstructured prose in description.

Consequence, stated plainly: the platform cannot currently answer "show me every trend across all 25 sectors that turns on HBM supply", or "on post-quantum cryptography", or "on NdFeB permanent magnets" — even though the research contains all three, in sectors 03, 04 and 13 respectively.

Specification for v1.0: a controlled vocabulary of roughly 200 technologies, each with a stable ID, a parent technology, a maturity anchor (research / pilot / production / commodity, aligned to the technical_maturity score) and a set of synonyms for extraction. Multi-valued. Populated by automated extraction against the vocabulary, with editorial confirmation required before a new technology enters the vocabulary (§3.6.1).

Worked example (target state): T-03-02 Advanced packaging, not wafers, is the binding constraint on AI computetechnologies: ["cowos","hbm","chiplet-2.5d"], which makes it joinable to T-01-06 and T-13-10 — two records in other sectors that turn on the same physical constraint and are currently unlinked (§3.5.1).

3.3.8 market_stage — single-valued, editorial

Allowed values (8): emerging_signal, early_adoption, growing, mainstream, mature, declining, reversing, speculative.

Distribution across 500 records (from the top-level stage field): early_adoption 122, growing 110, mainstream 68, emerging_signal 63, declining 55, mature 31, reversing 28, speculative 23.

Critical defect: the seed schema carries this facet twice — as the top-level stage and again as tags.market_stage. They agree on only 402 of 500 records (80.4%), and the copy inside tags has drifted to 12 values, admitting four non-canonical ones: emerging (21), cooling (3), transitional (1), current (1). The most common disagreements are mainstream/mature (15 records) and emerging_signal/emerging (13 records — pure vocabulary drift).

Fix: delete tags.market_stage. One field, one domain, CHECK constraint on the enum. Where the two disagree, stage is authoritative because it is the field the scoring model reads.

Worked example: T-13-02 Global industrial robot installations plateau near 540,000 unitsstage: "mature", direction: "steady". Installations flat, vendor share shifting — mature with internal reallocation, which is a different object from declining.

Note on the distinction from classification. classification (current / emerging_signal / cooling / overhyped, distributed 200/175/75/50) is an editorial judgement about the trend's epistemic status; stage is a claim about the market. They are orthogonal and must not be merged: T-24-19 is classification: overhyped with stage: mainstream — the narrative is mainstream, the substance is not.

3.3.9 time_horizon — single-valued, editorial

Allowed values (4): immediate_0_12m, near_1_3y, medium_3_7y, long_7y_plus.

Distribution: near_1_3y 240, immediate_0_12m 180, medium_3_7y 71, long_7y_plus 9.

Semantics: the horizon over which the trend's consequences land, not over which it was observed. A trend can be stage: mainstream and time_horizon: medium_3_7y if its effects are still working through — T-12-08 is exactly that.

The distribution is itself a finding. Only nine records in 500 — 1.8% — reach long_7y_plus, and one of them is T-13-19, the humanoid record classified overhyped. A trend-intelligence corpus that is 84% inside three years is an accurate reflection of what the evidence supports; it should be read as a limit on the product's forecasting reach, not as a claim that nothing matters after 2029.

Worked example: T-25-19 Agentic AI disintermediating online travel agenciesmedium_3_7y, against the sector's own precedent that "airlines will bypass the GDS" has been imminent since 2012 and Amadeus air bookings only turned negative (−3.7%) in H1 2026.

3.3.10 confidence — single-valued, computed with editorial override

Allowed values (3): high, medium, low. Distribution: 227 / 224 / 49.

Assignment rule (computable):

if verification_status == "triangulated"  and evidence_quality >= 4: high
elif verification_status == "single_source" or evidence_quality <= 2: low
else: medium

An editor may override downward freely and upward only with a recorded reason. In the seed database, confidence correlates with verification_status as designed: 347 records triangulated, 146 single_source, 7 unverified.

Worked example of the override mattering: T-13-19 carries confidence: high with classification: overhyped. Those are not in tension — we are highly confident that the category is overhyped, because the falsifying evidence is an S-4, a prospectus and the trade body's own statement. Confidence attaches to the record's claim, not to the trend's prospects.

3.3.11 direction — single-valued, editorial

Allowed values (4): accelerating, steady, decelerating, reversing. Distribution: 282 / 105 / 67 / 46.

Semantics: the second derivative of the trend, not the first. steady means "still happening at a constant rate", not "not happening". reversing is reserved for a change of sign in the underlying measure and requires a named indicator that turned.

Worked example: T-07-16 Property-catastrophe reinsurance pricing softens while the structural loss trend keeps risingdirection: "reversing", stage: "reversing", classification: "cooling". The named indicator: twelve consecutive quarters of cyber rate decline in the analogous line, and softening cat pricing against a rising loss trend. Without a named turning indicator, the correct value is decelerating.

3.3.12 impact — derived, computed

Impact is not a single editorial tag; it is the fifteen-dimension score block plus a derived band. Tagging impact directly invites inflation, which is why the seed schema decomposes it.

Components (integers 0–5, anchored): velocity, adoption, capital, revenue → momentum; breadth, depth, geographic_spread, customer_demand → reach; persistence, technical_maturity, strategic_importance → durability; regulatory_impact, social_impact → consequence; evidence_quality, source_diversity → the cap.

Derived band (impact_band), computed from composite_score:

Band Composite Records Example
defining ≥ 85 13 T-06-01 94.0 — incretin price/volume reset
major 70–84.9 149 T-01-01 75.8 — hyperscaler capex repricing
material 55–69.9 205 T-24-04 60.5 — AI upload flood, no listening
watch 40–54.9 119 T-04-19 50.4 — autonomous AI SOC
noise < 40 14 T-17-19 28.7 — creator-economy market sizing

The cap is the point. cap = 40 + 60 × mean(evidence_quality, source_diversity)/5 applies after the weighted composite, and 25 of 500 records are cap-bound — their measured momentum and reach were higher than their evidence would support, and the arithmetic, not an editor, held them down. T-17-06 (Consumer AI assistants as the fastest-monetising consumer app category on record) is the clearest instance: raw 83.5, cap 64.0, published 64.0.

3.3.13 source_type — single-valued per source, editorial at registration

Allowed values (27): regulatory_filing, government_data, company_ir, press_release, academic, patent, startup_db, vc_announcement, earnings, job_postings, app_store, search_trends, social, community, developer_activity, open_source, newsletter, trade_publication, conference, survey, market_report, procurement, trade_data, pricing, web_traffic, public_dataset, standards_body.

Reality across the 603 registered sources: 24 of 27 values used; one out-of-vocabulary value (trade_org, 1 source) that must be resolved to standards_body or trade_publication; and — the finding that matters for §7 — four allowed values have zero registered sources: patent, search_trends, social, community.

Value Sources Value Sources
government_data 130 startup_db 11
regulatory_filing 92 pricing 10
trade_publication 73 trade_data 6
company_ir 63 newsletter 4
market_report 58 developer_activity 3
earnings 30 procurement 3
survey 27 app_store 3
academic 25 web_traffic 2
press_release 24 vc_announcement 2
public_dataset 19 job_postings 1
standards_body 14 open_source 1
conference 1
patent, search_trends, social, community 0

Worked example: S-17-24 DSA Transparency Database, European Commission, source_type: public_dataset, access_method: bulk_download, cost: free, monitoring_priority: tier1_daily, common schema across TikTok, Meta, Snap, Pinterest and X. It is registered as a public_dataset rather than social because the type describes how the data is obtained and what its provenance is, not what it is about. That rule is what keeps the facet stable.

3.3.14 evidence_quality — single-valued, editorial, arithmetically enforced

Allowed values: integer 0–5, plus the companion verification_status enum (triangulated / single_source / unverified) and the per-source tier enum (A/B/C).

Anchors: 5 requires multiple Tier-A primary sources; Tier-C-only evidence caps at 0–1. source_diversity counts independent source organisations: 1 org = 1, 2 = 2, 3–4 = 3, 5–6 = 4, 7+ across at least two source types = 5.

Corpus state: 1,896 source citations across 500 records — 1,037 Tier A, 793 Tier B, 66 Tier C. 129 records cite Tier A exclusively; 60 records include at least one Tier C; no record is Tier-C-only, which is the contract's §0.2 rule holding in practice. Mean evidence_quality by classification: current 4.18, cooling 3.52, overhyped 3.38, emerging_signal 2.95 — the expected ordering, and the reason emerging signals score lower on composite by construction rather than by editorial preference.

Worked example of the facet doing its job: T-04-19 The autonomous AI SOC replacing Tier-1 analystsevidence_quality: 1, source_diversity: 1, verification_status: unverified, cap 52.0. Its five citations include two Tier-C aggregator posts, and a deliberate search returned no non-vendor efficacy evidence. The category has attracted over $300m in twelve months — Torq $140m at a $1.2bn valuation (Jan 2026), 7AI $130m (Dec 2025), Qevlar $30m (Mar 2026) — and the facet system still refuses to let it above 52.


3.4 The geography vocabulary

Geography is the facet where the seed database drifted furthest, and the drift is informative: analysts invented exactly the codes the specification failed to provide.

Observed: 65 distinct tokens across 1,625 assignments. They fall into five classes, three legitimate and two defective.

Class Examples observed Count Verdict
ISO-3166-1 alpha-2 country US 457, CN 96, JP 69, KR 43, IN 33 ~45 tokens correct
Supranational bloc EU 326, GCC 1 2 tokens correct, under-specified
Global scope Global 177 + GLOBAL 17 2 tokens correct, case-inconsistent
Non-ISO country alias UK 109 vs GB 20; CA-Canada 3 vs CA 27 3 tokens defective
Ad-hoc region Asia 14, APAC 13, LATAM 7 + LatAm 3 + LAC 1, Middle East 6, Africa 5, SEA 2, MEA 2, CIS 2, MENA 1, IMEA 1, Nordics 1, Caribbean 1 14 tokens defective, but the requirement is real

The canonical vocabulary, v1.0 — three tiers

Tier 1 — Country. ISO-3166-1 alpha-2, the full 249-code set, uppercase, no aliases. GB, not UK. CA is Canada and nothing else. A validator rejects any two-character token not in the ISO list.

Tier 2 — Supranational bloc. A closed, editorially governed list. Codes are three characters or more, except EU, which is safe because ISO-3166-1 exceptionally reserves it. This constraint is not cosmetic: the seed database uses AU 19 times meaning Australia, and a bloc vocabulary that admitted AU for the African Union would have silently corrupted 19 records.

Code Bloc Records Note
EU European Union 326 The second-most-used token in the corpus after US
NATO North Atlantic Treaty Organization 0 Specified, never used — see below
ASEAN Association of Southeast Asian Nations 0 Specified, never used; analysts wrote SEA instead
GCC Gulf Cooperation Council 1 T-08-06 only
EEA European Economic Area 0 added for MiCA/DSA scope precision
MERCOSUR Southern Common Market 0 added for LATAM trade records
OECD OECD 0 added for statistical-basis records

Tier 3 — Statistical region. This tier does not exist in the specification and the corpus proves it must. Fourteen ad-hoc tokens were invented across 59 assignments because there was nowhere else to put "the trend is Latin-America-wide and naming six countries would be false precision". Adopt UN M49 region and sub-region codes rather than inventing a third private vocabulary, with display aliases mapped on top:

Canonical (M49) Display alias Observed tokens it absorbs
M49-142 Asia Asia
M49-035 South-eastern Asia SEA
M49-419 Latin America and the Caribbean LATAM, LatAm, LAC
M49-029 Caribbean Caribbean
M49-002 Africa Africa
M49-145 Western Asia Middle East
M49-154 Northern Europe Nordics
REG-APAC Asia-Pacific APAC — a commercial region, not an M49 one; retained as a named exception
REG-MENA Middle East and North Africa MENA, MEA, IMEA
REG-CIS Commonwealth of Independent States CIS

Tier 0 — GLOBAL. One token, uppercase, meaning "no meaningful geographic concentration". 194 assignments once case is normalised, making it the third-most-used token. It is not a synonym for "we did not check" — T-24-19 carries geography: ["Global"] on a single-source record, which is a misuse: the Deezer statistic it rests on is one platform's, and the correct value would be the platform's markets. Validation rule: GLOBAL requires geographic_spread >= 4.

Cardinality: multi-valued, 1–9 observed, median 3. Mixing tiers in one array is permitted and common — T-18-01 carries ["US","CA","EU","APAC"], two countries, one bloc, one region — but a Tier-1 code that is contained by a Tier-2 or Tier-3 code in the same array must be intentional, i.e. the country is called out because it is doing something the bloc is not. The query layer expands blocs and regions to countries at read time, so geography contains "FR" matches a record tagged EU.

Assignment: automated named-entity recognition over the record's evidence and sources, proposing Tier 1 codes; editorial promotion to Tier 2 or Tier 3 where the mechanism is genuinely bloc- or region-wide. NER never assigns Tier 2 or 3 — the reason NATO and ASEAN have zero instances is that a bloc tag is an analytical claim about the scope of a mechanism, and no extractor should be making it.

Worked example: T-21-01 Stablecoin supply plateaus near $305bn["US","EU","Global","APAC","LATAM"]. Under v1.0 this normalises to ["US","EU","GLOBAL","REG-APAC","M49-419"] — and the validator flags it, correctly, because GLOBAL co-occurring with three narrower scopes is contradictory. The record should carry either GLOBAL or the specific scopes, not both.


3.5 Cross-cutting: what the corpus says the taxonomy is missing

Nine gaps, each measured against the file rather than asserted.

3.5.1 There is no cross-industry trend object, and the corpus needs one badly. related_trends contains 1,535 links. Zero of them cross an industry boundary. Not few — none. Meanwhile the same trend recurs across sectors under different record IDs: 33 records across 9 sectors turn on data-centre demand; 33 across 10 sectors on tariffs; 30 across 15 sectors reference app-store commission economics; 22 across 9 on agentic AI; 19 across 8 on AI training-data rights. The ranking document already identified the consequence — sector 24's analyst found that "AI training-data rights" was "a trend masquerading as a sector boundary" and recommended reversing the music/publishing merge on exactly those grounds.

Fix: a trend_cluster object with its own ID, a name, member trend IDs spanning sectors, and a cluster-level score. industry_id on the trend record stays single-valued (it is the record's home sector); membership in cross-sector clusters is many-to-many. This is the single highest-value schema change identified in Phase 1.

3.5.2 stage is stored twice and disagrees 19.6% of the time. §3.3.8. Delete tags.market_stage.

3.5.3 trend_type was never enum-constrained. 75 values for a 30-value vocabulary. §3.2. Enforce at write time, with the proposal queue of §3.6.1 as the escape valve.

3.5.4 companies[] and products[] are strings, not foreign keys. 16.1% exact-match rate against the entity table. §3.3.5.

3.5.5 The entity table is not deduplicated across sectors. 994 records, 924 distinct names, OpenAI ×9. Entity identity must be global; sector membership is an attribute, not a partition.

3.5.6 There is no technology facet at all. §3.3.7.

3.5.7 Geography has no region tier, and analysts invented one. §3.4.

3.5.8 Referential integrity is not enforced. 91 of 2,580 evidence items (3.5%) carry a source_idx that does not exist in their record's sources[] array — concentrated in sector 23, where 20 records account for most of the orphans. One related_trends reference is dangling (T-04-17T-04-22, which does not exist). Both are foreign-key constraints that a schema migration fixes permanently.

3.5.9 Two claim-type and entity-category values are out of vocabulary. claim_type: "survey" appears once against the six-value enum (it should be estimate); entity.category: "standards_body" appears once against the ten-value enum (it should be trade_org, or the enum should gain the value). Small, but they are the leading edge of exactly the drift that produced 75 trend types.


3.6 Governance

3.6.1 The five rules for changing a controlled vocabulary

  1. Values are never deleted and never redefined. A value is marked deprecated with a deprecated_at date and a successor pointer. Historical records keep their original value; the query layer follows the successor pointer, so a search for the new value finds old records and a search restricted to as_of < deprecated_at sees the corpus as it was. This is why startups and venture_capital are retired in §3.2 rather than removed: any future record carrying them resolves to business_models + industry_id: 11.
  2. Every vocabulary carries a version, and every record carries the version it was stamped under. taxonomy_version: "1.0" on the record. A record is not silently re-stamped when the vocabulary moves; it is re-stamped by an explicit, logged migration with a diff, or left alone. Cross-version comparison goes through a crosswalk table, not through hope. The 75 → 34 → 32 normalisation in §3.2 is the first such crosswalk and ships as data, not as prose.
  3. New values enter through a proposal queue, not through a write. A writer — human or agent — who needs a value that does not exist records the intended value in a proposed_value field alongside the nearest canonical value, which is what actually gets stored. The queue is reviewed on a fixed cadence. A proposed value is promoted when it crosses a usage threshold and an editor can state what query it enables that the nearest canonical value does not. narrative (14 uses, and it isolates three of the corpus's most instructive overhyped records) clears that bar; destination_management (1 use, indistinguishable from business_models) does not.
  4. Enum changes require a migration note in the changelog and a re-run of the validator over the full corpus. The validator is the artefact that makes rule 1 enforceable: it rejects out-of-vocabulary writes, reports drift, and is the thing that would have caught emerging vs emerging_signal on record 14 instead of record 500.
  5. Scores and their anchors are frozen harder than vocabularies. Changing what adoption: 3 means silently re-ranks 500 records. A scoring change is a major version bump, requires recomputation of the entire corpus, and both versions are retained.

3.6.2 Extension points, ranked by safety

Change Safety Procedure
Add a value to an open facet (technologies, products) safe proposal queue, no migration
Add a value to a closed facet (trend_type, source_type) moderate queue + threshold + editor rationale + validator re-run
Add a sector (26 = telecom) moderate new ID only; never reassign an existing one
Add a facet moderate nullable column; backfill is a separate, logged project
Deprecate a value moderate successor pointer mandatory
Change a facet's cardinality unsafe single → multi is a migration; multi → single is lossy and is refused
Redefine a value or a score anchor unsafe major version, full recompute, both versions retained

3.6.3 User-defined taxonomies (Phase 2)

The Phase 2 requirement is that a user can classify trends their own way — by their business units, their product lines, their watchlist themes. The failure mode is obvious and must be designed out: user vocabularies leaking into the canonical one and destroying cross-sector comparability, which is the platform's entire differentiator.

The overlay model.

  1. Separate namespace, separate storage. Canonical facets live in taxonomy.canonical.*. User facets live in taxonomy.custom.<workspace_id>.*. A custom facet can never share a name with a canonical one; the API rejects the collision at definition time rather than at query time.
  2. Custom tags are assignments, not edits. A custom classification is a row in a join table — (workspace_id, trend_id, custom_facet_id, value, assigned_by, assigned_at) — and never a mutation of the trend record. The canonical record is immutable to customers. This means a canonical re-score or re-tag propagates to every workspace without touching anyone's custom layer, and a workspace can be deleted without leaving residue.
  3. Custom taxonomies may reference canonical values but may not shadow them. A workspace may define "Our Q3 board themes" with rules like industry_id IN ("01","03") AND technologies CONTAINS "hbm". It may not define its own value named mainstream on market_stage. Derived-from-canonical facets are recomputed automatically when the canonical value changes; free-assignment facets are not.
  4. Custom values feed the proposal queue as evidence, anonymously and in aggregate. If forty workspaces independently define a tag that means "sovereign AI procurement", that is the strongest available signal that the canonical vocabulary is missing a value — and it is precisely the signal that would have surfaced market_structure and capital_markets before the corpus reached 500 records. Aggregate counts only; no workspace's taxonomy is visible to another.
  5. Export carries both layers, labelled. Any export marks each tag as canonical or custom:<workspace>, so a figure lifted from the platform into a board deck can be traced back to whether it was our classification or the customer's. This is the same discipline as the marketing claim label in §1.3: the provenance of a classification travels with it, always.

Companion documents: 10-mission-and-definitions.md (§1, the definitions this taxonomy operationalises), 11-industry-ranking.md (§2, the industry axis), 13-detection-methodology.md (§7, the signals that populate these facets), 01-research-contract.md (the schemas as issued to the research agents).

Research provenance
Source artifact
01-frameworks/12-taxonomy.md
Corpus date
15 September 2026
Prepared for this site
16 September 2026
Site publication
18 September 2026
Verification
Inherited; not fully rechecked