Make decisions when a critical field is missing without filling it with false precision or quietly excluding the case.
A blank field can mean zero, not applicable, not collected, not disclosed, inaccessible, suppressed, late or genuinely unknown. Combining those states changes conclusions. The UK Government Data Quality Framework advises producers to report missing records, missing important fields and reasons for absence, especially when missingness is systematic [1]. NIST's AI Risk Management Framework similarly says risks that cannot be measured should be documented rather than dropped [2]. Missing-data diligence therefore starts with classification and consequence: what is absent, through which mechanism, whose cases are more likely to be absent, and what decision becomes unsafe if the gap is ignored.
Use explicit absence codes
Replace empty cells with a controlled state: observed zero, not applicable, not asked, not reported, suppressed, source inaccessible, extraction failed, pending, or unknown. Preserve the raw value and the rule that assigned the state. A zero requires affirmative evidence that the measured event did not occur within the defined period. Silence in a filing, no API response or a dashboard dash does not meet that standard.
For each field, maintain an expectedness rule. Revenue may be expected for a public reporting segment but not for a private product line. Usage may be expected from internal telemetry but not a press release. Completeness is evaluated against expected records and fields, not against an ideal universe nobody can observe. The framework's warning about systematic missingness matters: if unsuccessful pilots stop reporting, the visible portfolio will overstate success even when every visible row is accurate [1].
Attach a consequence and next-best check
Every material gap receives a consequence statement. Specify whether it blocks a rate, lowers confidence, prevents segmentation, creates survivorship risk or leaves the direction unchanged. Then name the next-best check: an alternative official series, a later filing, a transaction log, a sampled manual review, a regulator docket, or a direct request routed through an authorized channel. 'Research further' is not a check because it has no target or stopping rule.
Rank checks by expected decision value, not ease. A cheap proxy that shares the missingness mechanism adds little. If private firms disclose only positive outcomes, another vendor case-study collection is not independent relief. If the authoritative page is inaccessible, record the access failure and look for an official bulk file, archived release or regulator copy. Do not circumvent controls. Stop when the likely decision change is smaller than the cost and delay of another check.
Run a sensitivity bound
Where possible, calculate the decision under plausible extreme assignments. Hypothetical example: 80 of 100 pilots report outcomes; 48 met the agreed threshold, 32 did not, and 20 are unknown. The observed success rate is 60% among reporters. The full-cohort rate is bounded between 48% if every unknown failed and 68% if every unknown succeeded. Reporting 60% as the portfolio rate silently assumes the missing group behaves like reporters.
Now connect the bound to action. If rollout requires at least 65%, the available data cannot clear the gate because only the optimistic bound does. The next-best check is to retrieve outcome logs for the 20 missing pilots or sample them under a documented rule. If rollout requires 45%, both bounds clear it, but missingness still matters for subgroup risk and future monitoring. NIST's framework supports recording unmeasured risks and repeatedly reassessing metrics as context changes [2].
- Code the reason for absence; never leave material blanks semantically empty.
- State the inferential consequence and the exact next-best check.
- Use worst- and best-case bounds before imputing a convenient middle.
Take it into the meeting
- Require affirmative evidence before encoding zero.
- Prioritize checks by their chance of changing the decision.
- Show whether plausible missing outcomes cross the action threshold.
Sources & boundaries
Source statements are attributed; the decision process is Signal Atlas analysis. Examples marked hypothetical are teaching inputs, not observed outcomes.
- Bounds can be wide and may not reveal the missingness mechanism.
- Authorized alternative sources may still share the same reporting bias.
- The protocol cannot manufacture evidence where no measurement or disclosure exists.
- The Government Data Quality Framework: guidanceUK Government · Source publication: not established · Retrieved 2026-09-19
Users should be told about missing records, missing fields and reasons for missingness. Systematic missingness may require adjustment in analysis.
- AI RMF CoreNational Institute of Standards and Technology · Source publication: not established · Retrieved 2026-09-19
Risks that cannot be measured should be documented. Measurement approaches and their effectiveness should be tracked and reassessed over time.