8 — The Trend-Scoring Framework
Phase 1 deliverable · research date 2026-09-15 · calibrated on 500 scored records
8.1 Design principle
The brief lists seventeen scoring inputs. Treating them as seventeen equal terms in one sum would be a mistake, because they are not the same kind of quantity. Fifteen are magnitude dimensions (how big, how fast, how far). Two — confidence and time horizon — are not magnitudes at all: confidence is a statement about the evidence, and time horizon is a facet you filter on, not a quantity you add.
Adding confidence to a magnitude sum produces the exact failure the brief warns against: a well-evidenced trivial trend and a poorly-evidenced enormous one converge on the same score. So the model separates them:
- 15 magnitude dimensions → grouped into four pillars → weighted sum →
raw_score - 2 evidence dimensions → an evidence factor that caps the result
- Time horizon → a filterable facet, never a score component
- Confidence → a published label derived from the evidence factor
8.2 The four pillars
Grouping is not cosmetic. Each pillar answers a different question, and users weight them differently by role: an investor cares most about Momentum, a policymaker about Consequence, a corporate strategist about Durability.
| Pillar | Weight | Question | Dimensions |
|---|---|---|---|
| Momentum | 30% | How fast is it moving, and is real money and usage behind it? | velocity, adoption, capital, revenue |
| Reach | 25% | How far does it extend? | breadth, depth, geographic_spread, customer_demand |
| Durability | 25% | Will it still be here in five years? | persistence, technical_maturity, strategic_importance |
| Consequence | 20% | Who does it hit, and does it attract rules? | regulatory_impact, social_impact |
8.3 The formula
M = mean(velocity, adoption, capital, revenue) / 5
R = mean(breadth, depth, geographic_spread, customer_demand) / 5
D = mean(persistence, technical_maturity, strategic_importance) / 5
C = mean(regulatory_impact, social_impact) / 5
raw = 100 × (0.30·M + 0.25·R + 0.25·D + 0.20·C)
E = mean(evidence_quality, source_diversity) / 5 # evidence factor, 0–1
cap = 40 + 60·E # evidence ceiling
score = min(raw, cap)
All fifteen inputs are integers 0–5 against a published rubric (§8.6). Every published
score shows its inputs, so any score is reconstructible from the record — which is also the
platform's principal defence against commercial capture (see 34-monetization.md).
8.4 The evidence cap, and an honest finding about it
The rule: a trend with no credible evidence (E=0) cannot score above 40. A trend needs E=1.0 — multiple Tier-A primary sources across at least two source types — to reach 100. This directly implements the brief's requirement that "a high volume of low-quality mentions must not outweigh a small number of strong, independent sources."
The finding we did not expect: across 500 records, the cap binds on only 25 (5%),
with a mean reduction of 5.1 points and a maximum of 19.5. And of those 25, 19 are
current trends and only 5 are emerging_signal — none are overhyped.
That is worth stating plainly rather than burying, because it means the cap is not what catches hype. Overhyped trends never reach the cap because they are already scoring low on the magnitude dimensions. What actually catches them is the Momentum pillar:
| Dimension | current | emerging | cooling | overhyped |
|---|---|---|---|---|
| adoption | 3.98 | 2.11 | 2.93 | 1.62 |
| revenue | 3.52 | 1.64 | 2.52 | 1.24 |
| customer_demand | 3.31 | 2.46 | 1.99 | 1.68 |
| capital | 2.92 | 2.09 | 1.77 | 2.44 |
The hype signature is arithmetically explicit: the overhyped cohort has the lowest
adoption, revenue and customer demand of any group — while carrying more capital than the
cooling cohort (2.44 vs 1.77). Money going in, demand not coming out. That divergence is
detectable automatically and should be a standing flag in Phase 2 (VC_HYPE, defined in
13-detection-methodology.md).
So the cap's real job is narrower than originally framed: it is a backstop against a
genuinely large trend being asserted on thin sourcing — which is precisely the 19
current records it caught. It is not the hype filter. Keeping both mechanisms is correct;
describing the cap as the anti-hype device would have been wrong.
8.5 Calibration on the seed corpus
Scores span 28.7 to 94.0, median 62.9 — a usable spread with no ceiling clustering.
| Classification | n | mean | median | range |
|---|---|---|---|---|
| current | 200 | 73.0 | 73.9 | 50.7 – 94.0 |
| cooling | 75 | 59.0 | 59.9 | 39.5 – 77.9 |
| emerging_signal | 175 | 56.9 | 57.0 | 36.2 – 85.2 |
| overhyped | 50 | 50.6 | 50.8 | 28.7 – 70.8 |
The ordering is the right one and the separations are meaningful. Two further checks:
coolingis distinguishable fromemergingdespite similar means, because the underlying profile is inverted: cooling trends carry high persistence (3.47) and technical maturity (3.88) with falling velocity (3.07) and weak demand (1.99) — mature things losing their market. Emerging trends carry high depth (3.73) with low adoption (2.11) and low persistence (2.12). A single score cannot tell these apart; the dimension profile can. This is the argument for always publishing the pillar breakdown alongside the composite.strategic_importancehas the highest mean (4.01) and lowest spread (sd 0.82), which means analysts used it least discriminatingly. Flagged for rubric tightening in Phase 2; a dimension that is nearly always 4 carries little information.
8.6 Rubric
Anchors are mandatory. 0 = absent or unknown; 5 = exceptional and documented.
Momentum — velocity: rate of change over 12 months; attention alone caps this at
3. adoption: deployed usage; no measurable users = 0–1, no exceptions. capital:
disclosed investment; rumoured rounds do not count. revenue: booked revenue attributable
to the trend; pre-revenue = 0.
Reach — breadth: subindustries touched. depth: how fundamentally it changes them
(cosmetic 1 → re-architecting 5). geographic_spread: 1 = one metro, 3 = one bloc,
5 = genuinely global. customer_demand: buyer pull as distinct from vendor push;
vendor-push-only = 1.
Durability — persistence: a trend under 12 months old cannot exceed 2. This is
the anti-fad dimension and the age rule is not discretionary. technical_maturity:
research 1 → pilot 2 → production 3 → commodity 5. strategic_importance: consequences for
who wins the sector.
Consequence — regulatory_impact: attached rule-making, in either direction.
social_impact: labour, health, safety, equity consequences.
Evidence — evidence_quality: 5 requires multiple Tier-A primary sources; Tier-C-only
caps at 1. source_diversity: count of genuinely independent source organisations —
1 org = 1, 2 = 2, 3–4 = 3, 5–6 = 4, 7+ across two or more source types = 5. Syndicated
copies of one wire story count once.
8.7 Interpretation bands
| Band | Reading | Action |
|---|---|---|
| 80–100 | Major, well-evidenced, already consequential | Act on it |
| 65–79 | Significant and established | Plan for it |
| 50–64 | Real but early, narrow, or contested | Monitor; revisit quarterly |
| 40–49 | Weak, or strong-but-poorly-evidenced | Investigate before acting |
| <40 | Evidence-capped or genuinely marginal | Do not act on this alone |
A score is never shown without its evidence factor. A 62 with E=0.95 and a 62 with E=0.45 are different objects: the first is a modest, well-understood trend; the second is possibly large and poorly understood. Displaying the composite alone would destroy that distinction, which is the whole product.
8.8 Mandatory score transparency
Every published score must show, per the brief: inputs (all 15 dimensions), source dates, missing data (dimensions scored 0 for absence of evidence versus genuine absence — these are different and must be distinguished in the UI), confidence, human editorial judgement applied, and whether the score is automated or manually reviewed.
Phase 1 seed data: 100% analyst-scored with human judgement, 69% triangulated,
55% Tier-A citations. Phase 2 introduces automated provisional scoring, and every
automated score must carry scored_by: automated until an editor reviews it. An
automated score may never be presented with the same visual weight as a reviewed one.
8.9 Known limitations
Stated because a scoring model that hides its weaknesses is worse than none.
- 0–5 integers are coarse. Two trends differing by 2 points are not meaningfully different. The bands in §8.7 are the honest resolution; the decimal is presentational.
- Analyst scoring varies between people. 25 analysts scored these records. Phase 2 needs calibration rounds with overlapping samples to measure and correct drift.
strategic_importanceis under-discriminating (§8.5) and needs a tighter rubric.- The weights are a judgement, not a derivation. They were chosen to reflect what decision-makers act on, and should be re-examined against user behaviour once there is any. The pillar structure lets users re-weight for their own role; the default is a default, not a truth.
- The model scores a trend's magnitude, not its relevance to you. A 90-scoring semiconductor trend may be irrelevant to a fashion buyer. Personalisation is a ranking layer on top, never a modification of the underlying score.
- Source artifact
- 01-frameworks/14-scoring-framework.md
- Corpus date
- 15 September 2026
- Prepared for this site
- 16 September 2026
- Site publication
- 18 September 2026
- Verification
- Inherited; not fully rechecked