Make an LLM talk about the real world without making things up
Your LLM describes something real, the weather, a location, live data, and lies with total confidence. Here's the discipline where the model only ever phrases facts a data layer actually asserts: every fact carries provenance, unsourced and stale claims are stripped, and when there's nothing solid it says nothing instead of inventing.
- System
- grounded generation
- Origin
- Abstracted from a production field-sensing engine.
Try this recipe in your own project
Read the files below first. The download for the whole kit asks for your email.
SKILL.md3.3 KBreference/grounding-gate.ts3.4 KB
Start in a test project with synthetic data. A skill file provides instructions; it does not grant permissions, start an autonomous process or establish compatibility with every agent host. Check this guide’s evidence and applicable file-specific licence.
The problem
Your feature has an LLM describe something real, the weather, a location, a user's data, live conditions, anything with a ground truth. And it lies with total confidence. It says it's clear when it's raining. It names a thing that isn't there. It states last week's number as today's. It fills a gap with a plausible invention rather than admitting it doesn't know. Anywhere the model touches real facts, you can't trust what it says, and you can't ship that.
What you get
A discipline where the model only ever phrases facts a data layer actually asserts, and everything else is caught or silenced:
- Every fact carries its provenance: source, timestamp, confidence, or it never reaches the model.
- The model phrases; it never asserts. It rewords what the data gave it; invented specifics get stripped before they reach the user.
- When there's nothing solid, it says nothing: silence is a real, first-class answer with a reason, not an error and not a fabrication.
- Stale facts don't leak: anything past its freshness window is dropped, not spoken as current.
- Accuracy fixes come from evidence, not anecdotes: so you stop chasing your tail on one-off complaints.
Use when / skip when
Use when an LLM speaks about facts with a ground truth a user could check, weather, place, prices, inventory, health, live data, anything real. The cost of a confident wrong statement is high.
Skip when the output is purely creative/subjective (there's no fact to get wrong), or the model only transforms text the user already provided and makes no claims about the world.
Ingredients
- A data/sensing layer that returns facts as structured values with
{ source, observedAt, confidence }, not prose. BYO, your APIs, your data. - An LLM used strictly as a phraser over those facts.
- A place to record per-claim outcomes for the replay loop (a table, a log).
The recipe
1. Provenance or it doesn't ship. Every fact the data layer emits carries
{ source, observedAt, confidence }. A value missing any of these is not passed to
the model. Provenance is a hard gate, not metadata.
2. Block confident claims that have no source. A claim asserting
strong confidence (for example 0.7 or above) with no source is structurally impossible. It's a
hallucination wearing a confidence score. Detect and drop it. See
reference/grounding-gate.ts.
3. Make silence a typed, first-class response. Absence is a shape, not a
failure: { silent: true, reason } (source-offline, stale, low-confidence,
thin-data, no-signal). The rendering layer must handle silence deliberately,
never invent to fill the hole, never 5xx.
4. Strip stale facts before the model sees them. Each fact declares a freshness window (TTL); the producer owns it. Anything older is removed from what the model is given. It can't speak what it never received. Conservative default TTL so unknown sources aren't spoken as fresh.
5. The model phrases, quality-gates catch leaks. Feed the model only sourced+fresh facts and instruct it to phrase, not add. Then run cheap deterministic gates on its output that catch the render layer undoing the discipline: raw null/placeholder strings, internal vocabulary leaking through, generic template fallbacks, and internal inconsistency (a field that contradicts another). Score the packet; a serious violation is a visible drop.
6. Degrade to silence, never fabricate. On timeout, thin data, or a source down, serve silence-with-reason for that axis, not stale data, not a guess. Wrap flaky sources in a circuit breaker (open after N consecutive failures, cool down, probe) so a struggling upstream fails to silence, fast.
7. Fix accuracy from evidence, not anecdotes (the replay loop). When someone
reports "this felt wrong," do not hand-patch from the anecdote, that way lies a
spiral of contradictory local fixes. Instead: snapshot the raw data behind every
generation; to investigate, replay that historical snapshot, grade it against your
current rules, and if a rule is new, grade the whole population to find every
similar case. Fix once, systematically. See reference/grounding-gate.ts for the
grade primitive.
8. Measure grounding, count silence as integrity. A single number:
grounding = (grounded + silent) / (grounded + silent + violations). Silence
counts toward integrity, not against it, a system that correctly says nothing is
being honest. Only violations (unsourced, stale, or confident without a source) score against you.
Evidence
From a production sensing engine running this discipline:
| Claim | Number |
|---|---|
| Confident claim with no source | blocked above a set cutoff, for example 0.7 |
| Freshness windows | set per kind of fact, for example minutes to hours for fast-moving readings and about a day for slow or computed ones, with a conservative default so unknowns aren't spoken as fresh |
| Circuit breaker | opens after a few consecutive failures, cools down for about a minute, then sends one probe |
| Grounding metric | (grounded + silent) / (grounded + silent + violations); silence is integrity |
| Quality gate | per-output score 0–100; one serious violation is a visible drop |
| Honesty over reach | when an upstream reports a value beyond what its sensor can measure, clamp it to the sensor's real ceiling. Under-claim rather than over-claim |
What might go wrong
- The anecdote spiral (the big one). Bites the moment you fix accuracy from individual complaints: each patch is a local hack, the patches contradict each other, and nobody ever checks whether the underlying signal is right. Months in, your rules are a pile of one-off fixes. Prevention is step 7, replay and grade against the population; let evidence drive fixes, never a single report.
- Confidence theater. A model emits a confidence score it didn't earn. Bites when confidence is decorative rather than gated to a source. The check in step 2 is the guard.
- The render layer undoes the grounding. The data layer is clean, then the final phrasing step invents or contradicts. Bites silently. The output quality gates (step 5) are there because the leak happens at the very end.
- Silence treated as failure. If "no data" throws an error or gets filled with a default, you've turned honesty into either an outage or a lie. Silence must be a designed, typed response.
- Tuning a threshold instead of demoting. When a claim class's confidence stops correlating with reality, the fix is not to nudge the number. It's to demote that class to silence until the signal earns its confidence back.
Keeping it healthy
- A calibration audit on a cadence: sample claims across contexts, check declared confidence against ground truth; a class that's meaningfully off its promise is a bug, and the fix is demotion to silence, not a threshold nudge.
- Every numeric threshold lives in a registry with where it came from and when it was last reviewed, no magic numbers buried in code.
- The replay archive keeps the raw data behind every generation, so any future rule can be graded against the past.
The skill
The skill and its reference code, in full. This is what an agent follows to build the pattern in your codebase. Copy any file here, or get every guide as one kit from naturate.io/recipes.
SKILL.md3.3 KB
---
name: grounded-generation
description: Make an LLM talk about the real world without making things up — data layer asserts facts with provenance, the model only phrases them, unsourced/stale claims are stripped, and it degrades to typed silence instead of inventing. Use when an LLM speaks about facts with a checkable ground truth (weather, place, prices, live data).
---
# Implement: grounded generation (no hallucinated facts)
You are stopping a developer's LLM feature from inventing real-world facts.
Implement the discipline below in their codebase. The core inversion: the model is
a PHRASER over facts a data layer asserts — it never originates a fact. Adapt to
their language, their data sources, and their model.
## Detect their setup
Where real facts enter (which APIs / data sources), where the model call is, and
whether facts currently reach the model as structured values or already baked into
a prose prompt (if baked-in, that's the first thing to fix — the model can't be
gated on provenance it can't see).
## Build this
1. **Structured facts with provenance.** Make the data layer return each fact as
`{ value, source, observedAt, confidence }`, not prose. A fact missing any of
source/observedAt/confidence does not pass to the model.
2. **The grounding gate** — classify every fact GROUNDED / SILENT / VIOLATION and
block violations before the model. Includes the cardinal-sin check (high
confidence + no source) and staleness (past its TTL). See
`reference/grounding-gate.ts`.
3. **Typed silence.** Absence is `{ silent: true, reason }`, handled deliberately
by the rendering layer — never invent to fill it, never error.
4. **Freshness.** Each fact declares a TTL (producer owns it); strip stale facts
before the model sees them; conservative default TTL for unknown sources.
5. **Model phrases only.** Prompt the model to reword the given facts and add
nothing. Then run cheap deterministic gates on its OUTPUT to catch leaks: raw
null/placeholder strings, internal vocabulary, generic template fallbacks, and
internal inconsistency (one field contradicting another). Score it.
6. **Degrade to silence.** Timeout / thin data / source down → silence-with-reason
for that axis. Wrap flaky sources in a circuit breaker (open after N failures,
cool down, probe).
7. **The replay loop.** Snapshot the raw facts behind every generation. When
accuracy is questioned, replay the snapshot and grade it against current rules
across the population — never hand-patch from one anecdote.
8. **A grounding metric.** `(grounded + silent) / (grounded + silent + violations)`
— silence counts as integrity.
## Guardrails to enforce
- No fact reaches the model without provenance; no high-confidence claim ships
without a source.
- Silence is a designed response, never an error or a filled-in default.
- Accuracy fixes are evidence-driven (replay + grade), never anecdote-driven.
- When a claim class stops correlating with reality, demote it to silence — don't
tune the threshold.
## Verify before done
Force a source offline and confirm the feature goes silent-with-reason (not stale,
not invented). Feed a fact with no source and confirm it's blocked. Confirm the
output gates catch an injected fabrication. Confirm a stale fact is stripped.
## Read alongside
`RECIPE.md` (the why + receipts) and `reference/grounding-gate.ts`.
reference/grounding-gate.ts3.4 KB
// Reference implementation — the grounding gate (steps 2, 4, 8) + the grade
// primitive (step 7). Provider-neutral. Adapt to your fact shape; don't copy blind.
//
// Core idea: classify every fact as GROUNDED, SILENT, or VIOLATION. Only GROUNDED
// facts reach the model. SILENT is honest absence. VIOLATION never ships and is
// what you measure against.
export interface Fact {
value: unknown;
source?: string;
observedAt?: number; // epoch ms
confidence?: number; // 0..1
ttlMinutes?: number; // producer-declared freshness; falls back to default
silentReason?: string; // set => a typed absence, not a claim
}
export type Verdict =
| { kind: 'GROUNDED'; fact: Fact }
| { kind: 'SILENT'; reason: string }
| { kind: 'VIOLATION'; reason: string; cardinal: boolean };
const CARDINAL_CONFIDENCE = 0.7; // confidence >= this with no source = impossible
const DEFAULT_TTL_MINUTES = 1440; // conservative: unknowns are NOT spoken as fresh
export function classify(fact: Fact, now = Date.now()): Verdict {
// Typed absence is first-class integrity, not a failure.
if (fact.silentReason || fact.value == null) {
return { kind: 'SILENT', reason: fact.silentReason ?? 'no-signal' };
}
const hasSource = !!fact.source;
// The cardinal sin: confident claim with no source.
if (fact.confidence != null && fact.confidence >= CARDINAL_CONFIDENCE && !hasSource) {
return { kind: 'VIOLATION', reason: 'confident claim with no source', cardinal: true };
}
// Provenance is mandatory for any real claim.
if (!hasSource || fact.observedAt == null) {
return { kind: 'VIOLATION', reason: 'missing provenance (source/observedAt)', cardinal: false };
}
// Staleness: past its freshness window => strip, don't speak as current.
const ttl = (fact.ttlMinutes ?? DEFAULT_TTL_MINUTES) * 60_000;
if (now - fact.observedAt > ttl) {
return { kind: 'VIOLATION', reason: 'stale beyond TTL', cardinal: false };
}
return { kind: 'GROUNDED', fact };
}
// Only GROUNDED facts are handed to the model. VIOLATIONs are dropped;
// SILENTs are passed to the render layer as typed absence to handle deliberately.
export function forModel(facts: Fact[], now = Date.now()): { grounded: Fact[]; silent: Verdict[] } {
const grounded: Fact[] = [];
const silent: Verdict[] = [];
for (const f of facts) {
const v = classify(f, now);
if (v.kind === 'GROUNDED') grounded.push(v.fact);
else if (v.kind === 'SILENT') silent.push(v);
// VIOLATION: intentionally dropped — never reaches the model or the user.
}
return { grounded, silent };
}
// The grounding metric — silence counts as integrity, only violations count against.
export function groundingRate(facts: Fact[], now = Date.now()): number {
let ok = 0, violation = 0;
for (const f of facts) {
const v = classify(f, now);
if (v.kind === 'VIOLATION') violation++;
else ok++; // GROUNDED and SILENT both count as integrity
}
const total = ok + violation;
return total === 0 ? 1 : ok / total;
}
// The grade primitive (step 7): re-run today's rules over a historical snapshot,
// so accuracy fixes are evidence-driven across the population, not anecdote-driven.
export function grade(snapshot: Fact[], now = Date.now()): { rate: number; violations: Verdict[] } {
const violations = snapshot
.map(f => classify(f, now))
.filter((v): v is Extract<Verdict, { kind: 'VIOLATION' }> => v.kind === 'VIOLATION');
return { rate: groundingRate(snapshot, now), violations };
}
Where this comes from
Abstracted from a production field-sensing engine whose whole purpose is to speak about the physical world without inventing it, its founding principle is that silence is a capability, not a fallback, and trust comes from restraint as much as accuracy. It shares patterns, with no source code.
Keep an AI feature reliable in production while you keep changing the prompt
You tweak the prompt and quietly break something else; the model makes things up; the provider hiccups and your feature goes dark. Here's how to run an AI feature so none of that happens, change prompts without a deploy, auto-check every output, never hard-fail, and prove a change is better before it ships.
Make an AI voice feature start playing in seconds, not after a long wait
You hand LLM-generated narration to a text-to-speech provider and users stare at a loading spinner for tens of seconds before anything plays, or you reach for a faster voice model and it drops your pauses, mispronounces markup, or wanders accent mid-clip. Here's the pipeline that streams audio starting in seconds, keeps pauses and loudness consistent across voice tiers, and degrades gracefully instead of hanging.