Keep an AI feature reliable in production while you keep changing the prompt
You tweak the prompt and quietly break something else; the model makes things up; the provider hiccups and your feature goes dark. Here's how to run an AI feature so none of that happens, change prompts without a deploy, auto-check every output, never hard-fail, and prove a change is better before it ships.
- System
- LLM prompt ops
- Origin
- Abstracted from a live consumer content-generation pipeline.
Try this recipe in your own project
Read the files below first. The download for the whole kit asks for your email.
SKILL.md5.2 KBreference/anti-patterns.ts2.1 KBreference/prompt-assembler.ts2.9 KB
Start in a test project with synthetic data. A skill file provides instructions; it does not grant permissions, start an autonomous process or establish compatibility with every agent host. Check this guide’s evidence and applicable file-specific licence.
The problem
You built a feature where an AI model writes something for your users, chat replies, summaries, product descriptions, support answers. It worked in the demo. Then in production: you tweak the prompt to fix one thing and quietly break another, and nobody notices for days. The model starts making things up. The morning the AI provider has a hiccup, your whole feature goes dark. And when you do change the prompt, you can't actually prove the new version is better than the old one, you just hope.
What you get
A way to run that AI feature so none of the above happens:
- Change the prompt without a code deploy: and without silently breaking it, because prompts live in one managed place with versions and one-click rollback, not buried in your code.
- Every single output gets quality-checked automatically: cheaply, in the background, so you catch the feature getting worse before your users do.
- The feature never hard-fails. When the model or the provider blips, it degrades to something usable instead of throwing an error at the user.
- Prove a prompt change is better before it ships: test the new version against real past inputs, compare the scores, then promote it.
Use when / skip when
Use when an LLM writes user-facing content on a live path and you will iterate on the prompt more than once a month. The whole apparatus exists so prompt iteration is safe and measurable.
Skip when the LLM output is internal-only, one-off, or human-reviewed before anyone sees it (an authoring tool with a human gate needs only steps 1–3 and 7; the runtime protections are overkill). Skip the judge layers when you have no quality bar written down yet, write the bar first.
Ingredients
- A prompt-management/observability service with versioned prompts, labels, traces, and scores (hosted or self-hosted). Use an account that belongs to the project it serves.
- An LLM provider with at least two model tiers (a frontier model for the flagship surface, a cheap fast model for judges and low-stakes surfaces).
- Env: two API keys for the prompt service, one for the provider. Nothing else.
- Expected running cost: judging adds a small fraction of the cost of the generation it checks.
The recipe
1. Make the prompt service the source of truth, and the repo the fallback.
Every prompt lives in the service under a {{feature-name}} key with a
production label. A copy of each prompt template stays in code, but it is a
cold-start fallback, nothing more. The decision that matters: runtime fetches
by label, humans promote by label, and the disk copy is never auto-pushed.
Editing the file in the repo must not change production.
2. Build one prompt assembler with a fail-open chain. All generation goes
through a single function: try the managed prompt → compile {{variables}}
into it → on any failure fall back to the in-code template → worst case return
something usable. The contract is that the user-facing request never fails
because the prompt service is down. Telemetry degrades; delivery doesn't.
3. Cache short, keep last-good forever. In-memory cache with a short TTL (for example 10 seconds, so a promotion is live within seconds) plus an indefinitely-held last-good version as the stale fallback. If you run a process cluster, broadcast invalidation cheaply (touching a file and checking its modified time per request is enough; no IPC needed).
4. Seed inert, promote by hand. Your prompt-seeding script's DEFAULT
behavior creates new versions under a seed-candidate label, never
production. Going live requires an explicit flag and a human promoting the
label in the service's UI. Your highest-value live prompt should be skipped by
the seeder entirely unless force-flagged. This step exists because of the first failure mode below. Keep a rollback script that can restore pre-regression
labels by timestamp.
5. Trace every generation and link it to the prompt version. One trace per call, carrying user/session ids, the audience/surface, and input. One generation object per model call, LINKED to the fetched prompt object, that link is what makes "which prompt version produced this" a column you can filter, and prompt-version A/B comparison possible at all.
6. Score with stable names, pushed fail-safe. Define a small vocabulary of score names ({{e.g. quality-pass (0|1), quality-findings (int), grounding-pass (0|1), user-feedback (-1|0|1)}}) and never rename them, dashboards and queries depend on stability. Score pushes are wrapped so a telemetry error can never reach the user path.
7. Layer the evals, deterministic first, judge second, grounding third:
- Deterministic detector (always, inline, <10ms). Regex/string rules encoding your named anti-patterns: hard rules (one hit = fail, and mechanically auto-repairable ones get repaired in-memory) and soft rules (vocabulary tics that only fail on accumulation). This is your written quality bar as code. It runs on 100% of output and costs nothing. {{YOUR_ANTI_PATTERNS, the named patterns from your own quality bar}}
- LLM-judge rubric (always, async, observe-only). A cheap model scores each output against {{N}} yes/no gates from your written rubric, fired after the response is sent, never gating delivery, failing open if the judge errors. Its job is to make quality a filterable score in your traces, not to block anything.
- Grounding judge (where facts matter, T≈0.1). Pass the verified facts plus the generated text to a cheap model asked one question: does the text contradict or invent? Bias it to pass, flag clear contradictions and invented specifics, allow paraphrase and poetic elaboration. Pair it with the source-asserts rule: the model may only phrase what your data layer asserts; strip or silence anything unsourced.
8. Tier models per surface, with fallback chains. Flagship surface gets the frontier model (off the critical path if you can pre-generate); everything else gets the cheap tier with the mid tier as robustness fallback; judges run on the cheap tier. Mark the system prompt for provider-side prompt caching and keep it universal (no per-user/per-place details in it) so the cache key holds across callers.
9. Close the loop with datasets from live traces. A script that pulls the last N traces (optionally filtered to failures via your pass/fail score) into a dataset in the service. That dataset is how a prompt change is tested against real inputs before its version gets the production label. Promotion without a dataset run is how regressions ship.
Evidence
From a content pipeline running this pattern in production across several surfaces.
| Claim | Number |
|---|---|
| Judge overhead per generation | a few percent of the generation's cost, 2–4s, async. Checking 100% of output is affordable |
| Provider prompt-cache win from a universal system prompt | ~50% input-token reduction on warm cache | | Promote-to-live delay under a short-TTL cache | seconds, no deploy | | Deterministic detector | <10ms, runs inline on everything | | LLM-rewrite-on-failure | tried, measured and disabled. It added 8–12s per miss and did no better than mechanical auto-repair plus a feed of recent outputs |
The last row is the useful one: the elaborate self-repair loop lost to a regex and a feed of recent outputs. Build the cheap layers first.
What might go wrong
- Your seed script regresses production. Bites when anything automated can write the production label, the classic trigger is a stale checkout whose on-disk prompts silently re-seed over a newer live version. Symptom: quality drops with no code deploy in sight, so nobody looks at prompts. Prevention is step 4's whole design: inert-by-default seeding, the flagship prompt excluded unless force-flagged, promotion as a separate human act, a label-rollback script written before you need it. If your seeder can touch production without a flag, this incident is scheduled, not hypothetical.
- Structured output truncates ugly. Bites when a tight max_tokens meets schema/tool-call mode: cutoff yields unparseable JSON and a hard failure, where free-text cutoff would have degraded into something still usable. Choose output mode per surface with cutoff behavior in mind, not just parsing convenience.
- A judge becomes an outage. Bites when any eval layer gates delivery: its latency, cost, and error rate are now on your user path. Every layer in this recipe is observe-only or fail-open by design; the one gating layer that was tried was measured and disabled (see Evidence).
- The fallback rots. Bites months in: the in-code fallback prompt drifts ever further behind the managed one, and the day the prompt service blips, users get last year's voice. The periodic fallback check under Keeping it healthy exists for this.
Keeping it healthy
What keeps this working after you build it:
- The rubric scores 100% of production output continuously; someone (or some agent) reads the score dashboard weekly and owns the question of whether quality dropped.
- Prompt promotion ritual: dataset run → compare scores → promote label. No promotion outside the ritual.
- The disk fallbacks drift stale by design (source of truth is the service); a periodic check that the fallback still produces acceptable output is part of the eval cadence, or your outage story quietly rots.
- Score names are frozen; adding is fine, renaming is a migration.
The skill
The skill and its reference code, in full. This is what an agent follows to build the pattern in your codebase. Copy any file here, or get every guide as one kit from naturate.io/recipes.
SKILL.md5.2 KB
---
name: llm-prompt-ops
description: Make an AI feature reliable in production while you keep changing the prompt — managed prompts with rollback, automatic quality checks on every output, graceful degradation when the provider blips, and prove-before-you-ship prompt changes. Use when an LLM writes user-facing content on a live path.
---
# Implement: reliable LLM feature in production
You are helping a developer add production reliability to a feature where an LLM
writes user-facing content (chat replies, summaries, generated text). Implement
the pattern below IN THEIR codebase, adapting to their language, framework, and
the prompt-management + observability service they use (or recommend one). Do
not copy any single vendor's SDK blindly — the pattern is provider-neutral; wire
it to whatever they have.
## Before you write code, detect their stack
Ask or infer: their language/runtime, their LLM provider(s), whether they use a
prompt/observability service (Langfuse, LangSmith, Braintrust, PromptLayer,
Helicone, or none), and where their generation call lives. Adapt every step to
what you find. If they have no prompt service, either wire the fallback layer
only (steps 2–3, 7) or recommend one — don't block.
## Build these, in this order
1. **Managed prompts, code fallback.** Move each prompt into the service under a
named key with a `production` label. Keep a copy in code as a COLD-START
fallback only. Rule: runtime fetches by label; humans promote by label; the
in-code copy is never auto-pushed. Editing the file must not change prod.
2. **One prompt assembler with a fail-open chain.** Route all generation through
a single function: fetch managed prompt → compile variables → on any failure
fall back to the in-code template → worst case return a usable default. It
MUST NOT throw on a service outage. See `reference/prompt-assembler.ts`.
3. **Cache short, keep last-good forever.** In-memory cache, short TTL (~10s), plus
an indefinitely-held last-good version as the stale fallback. Cluster? broadcast
invalidation cheaply (a touched file + mtime stat beats IPC).
4. **Seed inert, promote by hand.** The seeding script's DEFAULT creates versions
under a `candidate` label — never `production`. Going live needs an explicit
flag AND a human promoting in the UI. Keep a rollback script that restores a
prior label by timestamp. (This prevents the #1 real incident — a stale
checkout silently re-seeding over a newer live prompt.)
5. **Trace every generation, linked to the prompt version.** One trace per call
(user/session id, surface, input); one generation object linked to the fetched
prompt — that link is what makes "which prompt version produced this" filterable.
6. **Scores with stable names, pushed fail-safe.** A small frozen vocabulary
(e.g. `quality-pass` 0|1, `quality-findings` int, `grounding-pass` 0|1,
`user-feedback` -1|0|1). Wrap pushes so a telemetry error never reaches the user.
7. **Layer the evals — deterministic, then judge, then grounding:**
- **Deterministic detector** (always, inline, <10ms): regex/string rules for the
team's named anti-patterns. Hard rules fail (auto-repair the mechanical ones);
soft rules only fail on accumulation. See `reference/anti-patterns.ts`.
- **LLM-judge rubric** (always, async, observe-only): a cheap model scores each
output against N yes/no gates AFTER the response is sent. Never gates delivery;
fails open. Makes quality a filterable score.
- **Grounding judge** (where facts matter, temp ≈ 0.1): pass verified facts + the
text, ask "does it contradict or invent?", bias to pass. Pair with: the model
may only phrase what the data layer asserts; strip unsourced claims.
8. **Tier models per surface.** Flagship gets the frontier model (off the critical
path if you can pre-generate); everything else the cheap tier with a mid-tier
fallback; judges on the cheap tier. Mark the system prompt for provider-side
caching and keep it universal (no per-user details) so the cache key holds.
9. **Close the loop with datasets.** A script that pulls the last N traces
(optionally only failures, via the pass score) into a dataset. That dataset is
how a prompt change is tested against real inputs before it gets promoted.
## Guardrails to enforce as you build
- The user path never hard-fails on a prompt-service or judge error. Verify by
simulating an outage (bad key) — the feature must still respond.
- No eval layer gates delivery except the deterministic one's auto-repair.
- Score names are frozen; adding is fine, renaming is a migration.
- Don't build the LLM-rewrite-on-failure loop first (or at all early): where it was measured it added 8–12s per miss and did no better than auto-repair plus a feed of recent outputs. Ship the cheap layers; measure before adding the expensive one.
## Verify before done
Simulate a prompt-service outage (feature still responds), confirm a bad output
is caught by the detector, confirm a score lands on the trace linked to the
prompt version, and confirm promoting a new prompt version goes live without a
redeploy. Report what you wired and what you stubbed.
## Read alongside
`RECIPE.md` (the why + the honest receipts) and `reference/` (working code to
adapt, not copy verbatim).
reference/anti-patterns.ts2.1 KB
// Reference implementation — the deterministic anti-pattern detector (step 7a).
// Runs inline on 100% of output in <10ms. Hard rules fail (and the mechanical
// ones auto-repair); soft rules only fail on accumulation. Replace the RULES
// with YOUR team's named quality bar — these examples are illustrative, not law.
export interface Rule {
id: string;
test: RegExp;
repair?: (s: string) => string; // present => mechanically fixable, applied in-memory
}
// HARD rules: one hit = fail. Give the fixable ones a `repair`.
export const HARD_RULES: Rule[] = [
{ id: 'em-dash', test: /—/g, repair: s => s.replace(/\s*—\s*/g, ', ') },
{ id: 'double-space', test: / +/g, repair: s => s.replace(/ +/g, ' ') },
// ... your named anti-patterns go here (e.g. banned phrases, forbidden framings)
];
// SOFT rules: a single hit is fine; N+ within a window is a tic. One shared
// regex of "tell" words; the threshold is what flags it.
export const SOFT_TELLS = /\b(delve|nuanced|intricate|palpable|profound|evocative|tapestry|realm)\b/gi;
const SOFT_THRESHOLD = 2; // 2+ tells per ~200 words = flag
export interface CritiqueResult {
pass: boolean;
repaired: string; // the (possibly auto-repaired) text
hardFindings: string[]; // rule ids that fired after repair
softCount: number;
autoRepairs: number;
}
export function critique(input: string): CritiqueResult {
let text = input;
let autoRepairs = 0;
// Apply mechanical repairs first, count them.
for (const rule of HARD_RULES) {
if (rule.repair && rule.test.test(text)) {
const before = text;
text = rule.repair(text);
if (text !== before) autoRepairs++;
}
}
// Re-check hard rules AFTER repair — only unrepairable residue counts.
const hardFindings = HARD_RULES
.filter(r => { r.test.lastIndex = 0; return r.test.test(text); })
.map(r => r.id);
const softCount = (text.match(SOFT_TELLS) ?? []).length;
return {
pass: hardFindings.length === 0 && softCount < SOFT_THRESHOLD,
repaired: text,
hardFindings,
softCount,
autoRepairs,
};
}
reference/prompt-assembler.ts2.9 KB
// Reference implementation — the fail-open prompt assembler (step 2 + 3).
// Provider-neutral: you supply a PromptStore (Langfuse / LangSmith / Braintrust /
// your own) and an in-code fallback. Adapt types to your stack; do not copy blind.
//
// The one invariant: assemble() NEVER throws on a store outage. Telemetry may
// degrade; user delivery never fails.
export interface ManagedPrompt {
version: string;
template: string; // may contain {{variables}}
}
export interface PromptStore {
// Fetch the prompt labeled `production` for this name. May throw / time out —
// the assembler is responsible for surviving that.
fetch(name: string, label?: string): Promise<ManagedPrompt>;
}
// ---- short-TTL cache with an indefinitely-held last-good fallback (step 3) ----
interface CacheEntry { value: ManagedPrompt; fetchedAt: number; }
export class PromptCache {
private fresh = new Map<string, CacheEntry>();
private lastGood = new Map<string, ManagedPrompt>(); // never evicted
constructor(
private store: PromptStore,
private ttlMs = 10_000,
private now: () => number = () => Date.now(),
) {}
async get(name: string): Promise<ManagedPrompt | null> {
const hit = this.fresh.get(name);
if (hit && this.now() - hit.fetchedAt < this.ttlMs) return hit.value;
try {
const value = await this.store.fetch(name, 'production');
this.fresh.set(name, { value, fetchedAt: this.now() });
this.lastGood.set(name, value); // remember the last thing that worked
return value;
} catch {
// Store is down or slow. Serve the last thing that worked, forever.
return this.lastGood.get(name) ?? null;
}
}
}
// ---- the assembler: managed -> compile -> in-code fallback -> usable default ----
export interface AssembleArgs {
name: string;
variables?: Record<string, string>;
fallbackTemplate: string; // the in-code cold-start copy (step 1)
lastResort?: string; // returned if even compilation fails; default ''
}
export interface AssembledPrompt {
text: string;
source: 'managed' | 'fallback' | 'lastResort';
version?: string; // set when source === 'managed' (link this in traces, step 5)
}
function compile(template: string, vars: Record<string, string> = {}): string {
return template.replace(/\{\{\s*(\w+)\s*\}\}/g, (_, k) => vars[k] ?? '');
}
export async function assemble(
cache: PromptCache,
{ name, variables, fallbackTemplate, lastResort = '' }: AssembleArgs,
): Promise<AssembledPrompt> {
try {
const managed = await cache.get(name);
if (managed) {
return { text: compile(managed.template, variables), source: 'managed', version: managed.version };
}
} catch {
// fall through — never let the store break delivery
}
try {
return { text: compile(fallbackTemplate, variables), source: 'fallback' };
} catch {
return { text: lastResort, source: 'lastResort' };
}
}
Where this comes from
Abstracted from a live content-generation pipeline running every pattern here in production across multiple surfaces. It shares patterns, with no source code.
Build admin tools that humans and agents can share
Give people a clear review surface and agents a small set of typed actions. Keep permissions and business rules in one backend.
Make an LLM talk about the real world without making things up
Your LLM describes something real, the weather, a location, live data, and lies with total confidence. Here's the discipline where the model only ever phrases facts a data layer actually asserts: every fact carries provenance, unsourced and stale claims are stripped, and when there's nothing solid it says nothing instead of inventing.