Agent Recipesby Naturate

Make an LLM talk about the real world without making things up

Your LLM describes something real, the weather, a location, live data, and lies with total confidence. Here's the discipline where the model only ever phrases facts a data layer actually asserts: every fact carries provenance, unsourced and stale claims are stripped, and when there's nothing solid it says nothing instead of inventing.

In productionEvidence from one production system.
System
grounded generation
Origin
Abstracted from a production field-sensing engine.

Try this recipe in your own project

Read the files below first. The download for the whole kit asks for your email.

Get the kit (.zip)
  • SKILL.md
    3.3 KB
  • reference/grounding-gate.ts
    3.4 KB

Start in a test project with synthetic data. A skill file provides instructions; it does not grant permissions, start an autonomous process or establish compatibility with every agent host. Check this guide’s evidence and applicable file-specific licence.

The problem

Your feature has an LLM describe something real, the weather, a location, a user's data, live conditions, anything with a ground truth. And it lies with total confidence. It says it's clear when it's raining. It names a thing that isn't there. It states last week's number as today's. It fills a gap with a plausible invention rather than admitting it doesn't know. Anywhere the model touches real facts, you can't trust what it says, and you can't ship that.

What you get

A discipline where the model only ever phrases facts a data layer actually asserts, and everything else is caught or silenced:

  • Every fact carries its provenance: source, timestamp, confidence, or it never reaches the model.
  • The model phrases; it never asserts. It rewords what the data gave it; invented specifics get stripped before they reach the user.
  • When there's nothing solid, it says nothing: silence is a real, first-class answer with a reason, not an error and not a fabrication.
  • Stale facts don't leak: anything past its freshness window is dropped, not spoken as current.
  • Accuracy fixes come from evidence, not anecdotes: so you stop chasing your tail on one-off complaints.

Use when / skip when

Use when an LLM speaks about facts with a ground truth a user could check, weather, place, prices, inventory, health, live data, anything real. The cost of a confident wrong statement is high.

Skip when the output is purely creative/subjective (there's no fact to get wrong), or the model only transforms text the user already provided and makes no claims about the world.

Ingredients

  • A data/sensing layer that returns facts as structured values with { source, observedAt, confidence }, not prose. BYO, your APIs, your data.
  • An LLM used strictly as a phraser over those facts.
  • A place to record per-claim outcomes for the replay loop (a table, a log).

The recipe

1. Provenance or it doesn't ship. Every fact the data layer emits carries { source, observedAt, confidence }. A value missing any of these is not passed to the model. Provenance is a hard gate, not metadata.

2. Block confident claims that have no source. A claim asserting strong confidence (for example 0.7 or above) with no source is structurally impossible. It's a hallucination wearing a confidence score. Detect and drop it. See reference/grounding-gate.ts.

3. Make silence a typed, first-class response. Absence is a shape, not a failure: { silent: true, reason } (source-offline, stale, low-confidence, thin-data, no-signal). The rendering layer must handle silence deliberately, never invent to fill the hole, never 5xx.

4. Strip stale facts before the model sees them. Each fact declares a freshness window (TTL); the producer owns it. Anything older is removed from what the model is given. It can't speak what it never received. Conservative default TTL so unknown sources aren't spoken as fresh.

5. The model phrases, quality-gates catch leaks. Feed the model only sourced+fresh facts and instruct it to phrase, not add. Then run cheap deterministic gates on its output that catch the render layer undoing the discipline: raw null/placeholder strings, internal vocabulary leaking through, generic template fallbacks, and internal inconsistency (a field that contradicts another). Score the packet; a serious violation is a visible drop.

6. Degrade to silence, never fabricate. On timeout, thin data, or a source down, serve silence-with-reason for that axis, not stale data, not a guess. Wrap flaky sources in a circuit breaker (open after N consecutive failures, cool down, probe) so a struggling upstream fails to silence, fast.

7. Fix accuracy from evidence, not anecdotes (the replay loop). When someone reports "this felt wrong," do not hand-patch from the anecdote, that way lies a spiral of contradictory local fixes. Instead: snapshot the raw data behind every generation; to investigate, replay that historical snapshot, grade it against your current rules, and if a rule is new, grade the whole population to find every similar case. Fix once, systematically. See reference/grounding-gate.ts for the grade primitive.

8. Measure grounding, count silence as integrity. A single number: grounding = (grounded + silent) / (grounded + silent + violations). Silence counts toward integrity, not against it, a system that correctly says nothing is being honest. Only violations (unsourced, stale, or confident without a source) score against you.

Evidence

From a production sensing engine running this discipline:

ClaimNumber
Confident claim with no sourceblocked above a set cutoff, for example 0.7
Freshness windowsset per kind of fact, for example minutes to hours for fast-moving readings and about a day for slow or computed ones, with a conservative default so unknowns aren't spoken as fresh
Circuit breakeropens after a few consecutive failures, cools down for about a minute, then sends one probe
Grounding metric(grounded + silent) / (grounded + silent + violations); silence is integrity
Quality gateper-output score 0–100; one serious violation is a visible drop
Honesty over reachwhen an upstream reports a value beyond what its sensor can measure, clamp it to the sensor's real ceiling. Under-claim rather than over-claim

What might go wrong

  • The anecdote spiral (the big one). Bites the moment you fix accuracy from individual complaints: each patch is a local hack, the patches contradict each other, and nobody ever checks whether the underlying signal is right. Months in, your rules are a pile of one-off fixes. Prevention is step 7, replay and grade against the population; let evidence drive fixes, never a single report.
  • Confidence theater. A model emits a confidence score it didn't earn. Bites when confidence is decorative rather than gated to a source. The check in step 2 is the guard.
  • The render layer undoes the grounding. The data layer is clean, then the final phrasing step invents or contradicts. Bites silently. The output quality gates (step 5) are there because the leak happens at the very end.
  • Silence treated as failure. If "no data" throws an error or gets filled with a default, you've turned honesty into either an outage or a lie. Silence must be a designed, typed response.
  • Tuning a threshold instead of demoting. When a claim class's confidence stops correlating with reality, the fix is not to nudge the number. It's to demote that class to silence until the signal earns its confidence back.

Keeping it healthy

  • A calibration audit on a cadence: sample claims across contexts, check declared confidence against ground truth; a class that's meaningfully off its promise is a bug, and the fix is demotion to silence, not a threshold nudge.
  • Every numeric threshold lives in a registry with where it came from and when it was last reviewed, no magic numbers buried in code.
  • The replay archive keeps the raw data behind every generation, so any future rule can be graded against the past.

The skill

The skill and its reference code, in full. This is what an agent follows to build the pattern in your codebase. Copy any file here, or get every guide as one kit from naturate.io/recipes.

SKILL.md3.3 KB
---
name: grounded-generation
description: Make an LLM talk about the real world without making things up — data layer asserts facts with provenance, the model only phrases them, unsourced/stale claims are stripped, and it degrades to typed silence instead of inventing. Use when an LLM speaks about facts with a checkable ground truth (weather, place, prices, live data).
---

# Implement: grounded generation (no hallucinated facts)

You are stopping a developer's LLM feature from inventing real-world facts.
Implement the discipline below in their codebase. The core inversion: the model is
a PHRASER over facts a data layer asserts — it never originates a fact. Adapt to
their language, their data sources, and their model.

## Detect their setup

Where real facts enter (which APIs / data sources), where the model call is, and
whether facts currently reach the model as structured values or already baked into
a prose prompt (if baked-in, that's the first thing to fix — the model can't be
gated on provenance it can't see).

## Build this

1. **Structured facts with provenance.** Make the data layer return each fact as
   `{ value, source, observedAt, confidence }`, not prose. A fact missing any of
   source/observedAt/confidence does not pass to the model.

2. **The grounding gate** — classify every fact GROUNDED / SILENT / VIOLATION and
   block violations before the model. Includes the cardinal-sin check (high
   confidence + no source) and staleness (past its TTL). See
   `reference/grounding-gate.ts`.

3. **Typed silence.** Absence is `{ silent: true, reason }`, handled deliberately
   by the rendering layer — never invent to fill it, never error.

4. **Freshness.** Each fact declares a TTL (producer owns it); strip stale facts
   before the model sees them; conservative default TTL for unknown sources.

5. **Model phrases only.** Prompt the model to reword the given facts and add
   nothing. Then run cheap deterministic gates on its OUTPUT to catch leaks: raw
   null/placeholder strings, internal vocabulary, generic template fallbacks, and
   internal inconsistency (one field contradicting another). Score it.

6. **Degrade to silence.** Timeout / thin data / source down → silence-with-reason
   for that axis. Wrap flaky sources in a circuit breaker (open after N failures,
   cool down, probe).

7. **The replay loop.** Snapshot the raw facts behind every generation. When
   accuracy is questioned, replay the snapshot and grade it against current rules
   across the population — never hand-patch from one anecdote.

8. **A grounding metric.** `(grounded + silent) / (grounded + silent + violations)`
   — silence counts as integrity.

## Guardrails to enforce

- No fact reaches the model without provenance; no high-confidence claim ships
  without a source.
- Silence is a designed response, never an error or a filled-in default.
- Accuracy fixes are evidence-driven (replay + grade), never anecdote-driven.
- When a claim class stops correlating with reality, demote it to silence — don't
  tune the threshold.

## Verify before done

Force a source offline and confirm the feature goes silent-with-reason (not stale,
not invented). Feed a fact with no source and confirm it's blocked. Confirm the
output gates catch an injected fabrication. Confirm a stale fact is stripped.

## Read alongside

`RECIPE.md` (the why + receipts) and `reference/grounding-gate.ts`.
reference/grounding-gate.ts3.4 KB
// Reference implementation — the grounding gate (steps 2, 4, 8) + the grade
// primitive (step 7). Provider-neutral. Adapt to your fact shape; don't copy blind.
//
// Core idea: classify every fact as GROUNDED, SILENT, or VIOLATION. Only GROUNDED
// facts reach the model. SILENT is honest absence. VIOLATION never ships and is
// what you measure against.

export interface Fact {
  value: unknown;
  source?: string;
  observedAt?: number;   // epoch ms
  confidence?: number;   // 0..1
  ttlMinutes?: number;   // producer-declared freshness; falls back to default
  silentReason?: string; // set => a typed absence, not a claim
}

export type Verdict =
  | { kind: 'GROUNDED'; fact: Fact }
  | { kind: 'SILENT'; reason: string }
  | { kind: 'VIOLATION'; reason: string; cardinal: boolean };

const CARDINAL_CONFIDENCE = 0.7;   // confidence >= this with no source = impossible
const DEFAULT_TTL_MINUTES = 1440;  // conservative: unknowns are NOT spoken as fresh

export function classify(fact: Fact, now = Date.now()): Verdict {
  // Typed absence is first-class integrity, not a failure.
  if (fact.silentReason || fact.value == null) {
    return { kind: 'SILENT', reason: fact.silentReason ?? 'no-signal' };
  }
  const hasSource = !!fact.source;

  // The cardinal sin: confident claim with no source.
  if (fact.confidence != null && fact.confidence >= CARDINAL_CONFIDENCE && !hasSource) {
    return { kind: 'VIOLATION', reason: 'confident claim with no source', cardinal: true };
  }
  // Provenance is mandatory for any real claim.
  if (!hasSource || fact.observedAt == null) {
    return { kind: 'VIOLATION', reason: 'missing provenance (source/observedAt)', cardinal: false };
  }
  // Staleness: past its freshness window => strip, don't speak as current.
  const ttl = (fact.ttlMinutes ?? DEFAULT_TTL_MINUTES) * 60_000;
  if (now - fact.observedAt > ttl) {
    return { kind: 'VIOLATION', reason: 'stale beyond TTL', cardinal: false };
  }
  return { kind: 'GROUNDED', fact };
}

// Only GROUNDED facts are handed to the model. VIOLATIONs are dropped;
// SILENTs are passed to the render layer as typed absence to handle deliberately.
export function forModel(facts: Fact[], now = Date.now()): { grounded: Fact[]; silent: Verdict[] } {
  const grounded: Fact[] = [];
  const silent: Verdict[] = [];
  for (const f of facts) {
    const v = classify(f, now);
    if (v.kind === 'GROUNDED') grounded.push(v.fact);
    else if (v.kind === 'SILENT') silent.push(v);
    // VIOLATION: intentionally dropped — never reaches the model or the user.
  }
  return { grounded, silent };
}

// The grounding metric — silence counts as integrity, only violations count against.
export function groundingRate(facts: Fact[], now = Date.now()): number {
  let ok = 0, violation = 0;
  for (const f of facts) {
    const v = classify(f, now);
    if (v.kind === 'VIOLATION') violation++;
    else ok++; // GROUNDED and SILENT both count as integrity
  }
  const total = ok + violation;
  return total === 0 ? 1 : ok / total;
}

// The grade primitive (step 7): re-run today's rules over a historical snapshot,
// so accuracy fixes are evidence-driven across the population, not anecdote-driven.
export function grade(snapshot: Fact[], now = Date.now()): { rate: number; violations: Verdict[] } {
  const violations = snapshot
    .map(f => classify(f, now))
    .filter((v): v is Extract<Verdict, { kind: 'VIOLATION' }> => v.kind === 'VIOLATION');
  return { rate: groundingRate(snapshot, now), violations };
}

Where this comes from

Abstracted from a production field-sensing engine whose whole purpose is to speak about the physical world without inventing it, its founding principle is that silence is a capability, not a fallback, and trust comes from restraint as much as accuracy. It shares patterns, with no source code.

On this page