Agent Recipesby Naturate

Make an AI voice feature start playing in seconds, not after a long wait

You hand LLM-generated narration to a text-to-speech provider and users stare at a loading spinner for tens of seconds before anything plays, or you reach for a faster voice model and it drops your pauses, mispronounces markup, or wanders accent mid-clip. Here's the pipeline that streams audio starting in seconds, keeps pauses and loudness consistent across voice tiers, and degrades gracefully instead of hanging.

In productionEvidence from one production system.
System
voice AI pipeline (progressive/streaming TTS)
Origin
Abstracted from a production voice pipeline that streams AI-authored narration as audio.

Try this recipe in your own project

Read the files below first. The download for the whole kit asks for your email.

Get the kit (.zip)
  • SKILL.md
    6.1 KB
  • reference/chunk-and-pace.ts
    7.5 KB
  • reference/failure-ladder.md
    2.1 KB

Start in a test project with synthetic data. A skill file provides instructions; it does not grant permissions, start an autonomous process or establish compatibility with every agent host. Check this guide’s evidence and applicable file-specific licence.

The problem

You generate narration with an LLM, a guided script, a narrated report, a voice reply, and hand the text to a text-to-speech provider. In the obvious implementation, the user stares at a loading spinner for tens of seconds before anything plays, because you wait for the whole piece to synthesize before you send back a single byte. Reach for a faster voice model and it drops your carefully authored pauses, mispronounces cue markup, or wanders accent mid-clip. And the failure modes are invisible until real users hit them: one slow request quietly holds a shared capacity slot and starves every other concurrent user, a provider hiccup throws instead of degrading, and the finished audio's loudness drifts from one generation to the next.

What you get

  • Users hear the first words in seconds: not tens of seconds, audio starts arriving before the rest of the piece has even finished synthesizing.
  • A speed/quality dial per surface: a fast, cheap voice tier for the moment someone is actively waiting, and a slower, higher-quality tier for anything you can prepare ahead of time.
  • Authored pauses survive the trip through the model: faithfully, no matter which voice tier is doing the talking.
  • Every finished piece has the same perceived loudness: regardless of voice, model tier, or day it was generated.
  • The pipeline degrades gracefully to a slower, already-proven path when the fast path breaks, never a spinner that never resolves, never a hard error.
  • One runaway request can't starve everyone else sharing the provider's capacity.

Use when / skip when

Use when an LLM or any dynamically generated script needs to become spoken audio on a live, latency-sensitive path, and you'll run this at real scale (dozens-plus concurrent generations against a shared provider quota).

Skip when audio is generated once, ahead of time, with nobody waiting live (a nightly batch render), you only need the loudness and concatenation pieces (steps 6–7), not the streaming apparatus. Skip entirely if TTS calls are rare enough that a shared concurrency ceiling will never bind.

Ingredients

  • An LLM that supports streaming output (token/SSE deltas).
  • A TTS provider offering a fast/cheap tier and a slower/higher-quality tier, a documented per-account concurrency ceiling, and ideally a low-latency streaming-input mode.
  • ffmpeg (or an equivalent) on the audio-serving host, for silence synthesis, concatenation, and loudness normalization.
  • Object storage with presigned URLs for delivering chunks/files to a client.
  • A feature flag / percentage-rollout mechanism.

The recipe

1. Pick your delivery shape before you touch a provider. Three shapes exist: plain batch (generate everything, upload once, hand back one URL); progressive multi-track (return after the first chunk, keep synthesizing and uploading the rest in the background, the player drains chunks gaplessly); and a true bytestream (proxy the provider's own audio stream straight to the client). Pick progressive if your client already has, or can cheaply build, a gapless multi-track queue: it captures most of the latency win without a client audio-engine rewrite, and it degrades to plain batch trivially.

2. Split the script into synthesis chunks at its natural seams, not by size alone. Break at author-placed pause markers first. Both extremes are wrong: too few chunks and one slow chunk (usually the last) sits on the critical path of the whole request; too many chunks and the tail becomes a queue of stragglers that blows your latency budget once real per-request round-trip overhead dominates. Tune target chunk size empirically against the specific tier's measured per-request latency, a size that's right for a fast tier is wrong for a slow one.

3. Preserve authored pauses as real silence, translated per model tier. A pause marker in your source script must survive to a real gap in the audio. Expressive/high-quality model tiers often only honor inline pause cues (a "[short pause]"/"[long pause]"-style tag) and will otherwise run straight through or mis-render a literal pause syntax; fast/cheap tiers often honor literal pause syntax natively. Classify pauses by depth: small in-line pauses can ride as whatever native syntax that specific tier understands; the deepest, author-marked holds should become real reinjected silence at a chunk boundary, synthesize a silent clip and splice it in, so a model that doesn't understand your pause syntax can never swallow it. See reference/chunk-and-pace.ts.

4. Tier your voice model to the surface, and respect the concurrency ceiling as the real constraint. Give the at-tap/interactive path (someone actively waiting) the fast/cheap tier; reserve the expensive, most expressive tier for anything you can pre-generate ahead of need, so its slower per-request cost never sits on a human's critical path. The provider's concurrency ceiling is usually account-wide, shared across every concurrent user and any background pre-generation, so treat it as a semaphore your dispatcher must respect, and make a rate-limit response retry with backoff rather than fail the whole generation; a single 429 shouldn't dump an entire piece to the fallback tier.

5. Give every synthesis request its own strongly-referenced timeout, chained to a caller-level abort. A composite/derived abort signal built from a caller signal plus a timeout, with nothing holding the timeout strongly, can silently fail to fire in some runtimes, the request just hangs. Use an explicit timer you hold directly and clear in a finally, so a hung request always actually dies and frees its concurrency slot. One request stuck open starves every other concurrent user of that shared quota.

6. Concatenate by re-encoding, never by raw byte-copy. Splicing compressed audio chunks (and any reinjected silence) together with a raw byte-copy is fast, but it leaves either wrong duration metadata (a player's progress bar reads only the first chunk's header) or audible micro-glitches at seams where frame boundaries don't align. Re-encode through the format's real encoder in the same pass, still well under a second for a multi-minute piece, and you get one continuous, gapless file with correct metadata.

7. Normalize loudness in the same encoding pass. Fold a loudness-normalization filter into step 6's re-encode (one filter pass, no added latency) targeting a standard loudness level with a true-peak ceiling, so every generated piece sounds the same volume regardless of voice, model tier, or generation date, instead of drifting piece to piece.

8. Build the failure ladder before you ship, and detect every rung. Name every point that can fail, upstream text generation stalls, TTS fails to connect, TTS returns no audio within budget, a mid-stream error, and give each a wall-clock detection budget and a fallback action. The last rung is always a plain batch/fallback path that already works, never new code written just for the failure case. A tier that's higher quality in the demo but empirically breaks your latency budget on the interactive path isn't a bug to patch. It's a signal to keep it pre-generate-only until proven otherwise.

9. Instrument the seams, not just the total. Log time-to-first-audio-byte, per-chunk synthesis time, total stream duration, and which fallback rung (if any) fired, per generation. A single "it took 20 seconds" number can't tell you whether the slow part was upstream generation, the TTS round-trip, or concatenation, and you need that breakdown to tune step 2's chunk size against reality instead of guessing.

Evidence

From a live voice pipeline turning AI-generated narration into streamed audio.

ClaimNumber
Progressive delivery's first-audio winfirst chunk observed ready in ~2–3s in a real smoke test, vs a ~20–60s whole-piece batch bake
Fast/cheap tier vs expressive tierroughly 5–6x faster, about half the per-request cost, at a real quality tradeoff judged acceptable for the interactive path
Expressive tier, cold, plain-concat modelseveral-tens-of-seconds floor even before variance, whichever chunk finishes last always sits on the critical path
Provider concurrency ceilingaccount-wide, shared across concurrent users and background pre-generation, a small single-digit number of simultaneous requests at the plan tier in use
Re-encode cost at concat timelow hundreds of milliseconds for a multi-minute piece, cheap insurance against seam glitches and wrong duration metadata
A real slice of the fallback rate, pre-fixroughly a quarter of all fallbacks traced to one bug class: a chunk containing only non-spoken cue markup synthesized as literally empty text, which the provider rejected outright

What might go wrong

  • The silent-timeout footgun. Bites when a composite/derived abort signal is built from a caller signal plus a timeout, but nothing keeps that timeout strongly referenced, some runtimes garbage-collect the timer and the abort never fires. Symptom: a single request hangs for minutes, invisibly holding a shared concurrency slot, and every other concurrent generation queues behind it or fails. Prevention: an explicit timer object you hold and clear yourself, never a derived signal you don't retain.
  • The pause-translation mismatch. Bites when you swap voice model tiers (or A/B two) without re-checking how each honors your pause syntax, one tier honors literal pause markup natively, another needs inline cue tags instead. Symptom: authored silences "run straight through," and a piece with deliberate holds reads flat and rushed. Prevention: treat pause-translation as a per-tier property you test explicitly, never assume identical pacing across tiers.
  • Accent/persona drift. Bites when an expressive model synthesizes many small chunks of the same "voice" independently, with no cross-chunk context, at a low stability/consistency setting, each chunk restarts its interpretation of the voice and can wander mid-piece. Prevention: raise the model's stability setting for any chunked-independent-synthesis path; accept the small expressiveness cost.
  • The byte-copy seam. Bites the first time someone "optimizes" concatenation to skip re-encoding for speed. Symptom: intermittent audible micro-glitches at chunk boundaries, and a player's duration/progress reads only the first chunk's metadata so playback position visibly desyncs from real audio length. Prevention: step 6, always.
  • The over-chunked tail. Bites when chunk size is tuned by word count alone without accounting for real per-request round-trip cost on the slower tier, many small chunks create a long tail of late-dispatched requests contending for the same concurrency slots, and on the tier where per-request cost is high, the straggler tail blows the latency budget and the whole generation falls back to the faster tier. Prevention: tune target chunk size against measured per-request latency on the specific tier in use, not a fixed word count right for a different tier.
  • Buffering that defeats streaming. Bites when a compression middleware or edge cache sits in front of the streaming response and buffers the whole body before forwarding, the client sees zero benefit from a carefully built streaming pipeline because nothing arrives until the buffer flushes. Prevention: explicitly disable compression and any edge-buffering on the streaming route; verify with a raw client that bytes arrive incrementally, don't trust that it streamed in one browser tab.

Keeping it healthy

  • A synthetic per-generation trace (time-to-first-byte, per-chunk timing, which fallback rung fired) feeding a dashboard someone actually reads, a silent creep in fallback rate is the leading indicator that a provider quietly changed model behavior.
  • A percentage-rollout flag on the streaming/progressive path, keyed per-user so rollout is safe to move up or roll back without a redeploy, and so one user's experience stays consistent session to session.
  • Periodic listening spot-checks against the loudness target and pause fidelity, a provider can silently change model behavior on their end; only ears (or an automated loudness scan) catch a drift a latency dashboard won't.
  • Whoever owns this pipeline re-reads the failure ladder (step 8) after every provider model/tier change, the fallback rungs and their budgets are the part most likely to silently rot as providers ship new versions.

The skill

The skill and its reference code, in full. This is what an agent follows to build the pattern in your codebase. Copy any file here, or get every guide as one kit from naturate.io/recipes.

SKILL.md6.1 KB
---
name: voice-ai-pipeline
description: Turn AI-generated text into voice audio that starts playing in seconds instead of making users wait for the whole file — chunked/progressive TTS delivery, pause fidelity across voice model tiers, loudness normalization, and a failure ladder that degrades to a proven batch fallback instead of erroring. Use when an LLM or dynamic script needs to become spoken audio on a live, latency-sensitive path.
---

# Implement: a streaming/progressive text-to-speech pipeline

You are helping a developer turn AI-generated (or any dynamic) text into voice
audio that starts playing quickly, instead of the user waiting for the whole
piece to render. Implement the pattern below IN THEIR codebase, adapting to
their language, framework, TTS provider, and audio player. Do not copy any
single vendor's SDK blindly — the pattern is provider-neutral; wire it to
whatever they have.

## Before you write code, detect their stack

Ask or infer: their language/runtime, their TTS provider (ElevenLabs, Azure
Speech, Google Cloud TTS, AWS Polly, OpenAI TTS, or none yet), whether that
provider offers both a fast/cheap tier and a higher-quality tier, its
documented concurrency limits, and what their client-side audio player can
already do (does it support gapless multi-track queueing, or only a single
file?). Adapt every step to what you find. If they have no player capable of
multi-track playback, either recommend the smallest viable upgrade or scope
down to the plain-batch path (skip steps 1–5, do 6–9 only).

## Build these, in this order

1. **Pick the delivery shape.** Default to progressive multi-track delivery
   if the client can (or can cheaply be made to) drain a queue of audio
   chunks gaplessly: return after the first chunk is ready, keep
   synthesizing + uploading the rest in the background. If the client can
   only play one file, do plain batch (skip to step 6) — don't force a
   client rewrite to unlock this recipe.

2. **Split text into synthesis chunks at natural seams.** Break at
   author-placed pause markers first, not by raw word count. Tune the target
   chunk size against the SPECIFIC tier's measured per-request latency: too
   few chunks puts one slow chunk on the critical path; too many creates a
   straggler tail that blows budget on a slow tier. See
   `reference/chunk-and-pace.ts` for the chunking + pause-classification
   shape.

3. **Translate pauses per model tier, never assume one syntax works
   everywhere.** Classify each pause as shallow (ride inline, using whatever
   native pause syntax that tier honors) or deep/authored (make it REAL
   reinjected silence spliced in at a chunk boundary — synthesize a short
   silent clip and stitch it in). Verify by ear on EACH tier you support; do
   not assume a script that paces correctly on one tier paces the same on
   another.

4. **Tier the voice model to the surface.** Fast/cheap tier for the
   interactive, someone's-waiting path. Slower/higher-quality tier ONLY for
   anything pre-generated ahead of need (a cron, a "prepare while the user
   is still browsing" trigger). Respect the provider's concurrency ceiling
   as a semaphore shared across ALL concurrent work (live + background); on
   a rate-limit response, retry with backoff instead of failing the whole
   generation.

5. **Give every synthesis call an explicit, strongly-held timeout chained to
   a caller abort.** Do not rely on a composite/derived AbortSignal alone —
   in some runtimes an unreferenced derived timer can be garbage-collected
   and silently never fire, hanging the request and starving a shared
   concurrency slot. Hold your own timer, clear it in a `finally`.

6. **Concatenate by re-encoding, never raw byte-copy.** Byte-copy concat is
   faster but leaves audible seam glitches (frame-boundary misalignment) and
   wrong duration metadata (a player reads only the first chunk's header).
   Re-encode through the format's real encoder in the same pass — still
   sub-second for a multi-minute piece.

7. **Fold loudness normalization into the same re-encode pass.** One filter
   pass, no added latency, targeting a standard loudness level with a
   true-peak ceiling so output volume doesn't drift piece to piece. See
   `reference/chunk-and-pace.ts` for where this slots into the concat step.

8. **Build the failure ladder, and make every rung detectable.** Name each
   failure point (upstream generation stalls, TTS connect fails, TTS returns
   no audio within budget, mid-stream error) with a wall-clock detection
   budget and a fallback action. The last rung is always a plain,
   already-proven batch path — never new code written only for the failure
   case. See `reference/failure-ladder.md` for the table shape.

9. **Instrument the seams.** Log time-to-first-audio-byte, per-chunk
   synthesis time, total stream duration, and which fallback rung (if any)
   fired — per generation, not just a single total.

## Guardrails to enforce as you build

- The interactive path never uses the expensive/slow tier live — only
  pre-generated ahead of need.
- A single hung synthesis request cannot hold a shared concurrency slot
  forever — verify the timeout actually fires under a simulated hang.
- Concatenation always re-encodes; never leave a fast byte-copy path in the
  code even "temporarily."
- Every fallback rung actually degrades to something that works — verify by
  forcing each failure point (bad API key, injected timeout, truncated
  response) and confirming the user still gets audio.
- Compression middleware / edge caching is explicitly disabled on the
  streaming route — verify bytes arrive incrementally with a raw client, not
  just "it played in a browser tab."

## Verify before done

Simulate each rung of the failure ladder (upstream stall, TTS connect
failure, TTS timeout, mid-stream error) and confirm the user still gets
audio via the batch fallback. Confirm a deep authored pause survives
playback on every voice tier in use. Confirm two consecutive generations of
the same script land at the same perceived loudness. Report what you wired
and what you stubbed.

## Read alongside

`RECIPE.md` (the why + the honest receipts) and `reference/` (working code
to adapt, not copy verbatim).
reference/chunk-and-pace.ts7.5 KB
/**
 * reference/chunk-and-pace.ts — provider-neutral shape for splitting a
 * script into synthesis chunks, classifying pauses, and re-assembling the
 * final audio with correct pacing and consistent loudness.
 *
 * This is a pattern to ADAPT, not a drop-in library. Swap the TTS call, the
 * silence-synthesis call, and the concat/loudnorm call for your own
 * provider's SDK / your own audio toolchain. No vendor-specific code here —
 * see RECIPE.md steps 2-3 and 6-7 for the reasoning behind each shape.
 */

// ── 1. Parse the script into text blocks + standalone pause markers ───────

type ScriptItem =
  | { type: 'text'; text: string }
  | { type: 'pause'; seconds: number }

/**
 * Split a script on blank lines into ordered items: spoken text blocks and
 * standalone `<pause seconds="N"/>`-style markers the author placed between
 * passages. Adapt the marker syntax to whatever your authoring pipeline
 * uses (SSML <break/>, a custom tag, plain "..." conventions).
 */
function parseScript(script: string): ScriptItem[] {
  const blocks = script.split(/\n\s*\n/).map((b) => b.trim()).filter(Boolean)
  return blocks.map((b) => {
    const m = b.match(/^<pause seconds="(\d+(?:\.\d+)?)"\s*\/?>$/i)
    return m ? { type: 'pause', seconds: parseFloat(m[1]) } : { type: 'text', text: b }
  })
}

// ── 2. Group into chunks, deciding which pauses become REAL silence ───────

interface Chunk {
  rawText: string
  /** Seconds of REAL reinjected silence to splice in after this chunk. */
  silenceAfterSec: number
}

/**
 * Group text blocks into at most `maxChunks` synthesis chunks, choosing the
 * `maxChunks - 1` LARGEST standalone pauses as real chunk-boundary silence.
 * Every other pause (smaller, or falling inside a merged chunk) gets
 * re-inlined as a token for the per-tier intra-chunk transform to handle
 * natively (see step 3 below) — it still paces, just not as spliced silence.
 *
 * Tune `maxChunks` and your own target-words-per-chunk against the SPECIFIC
 * tier's measured per-request latency (RECIPE.md step 2) — a size that's
 * right for a fast tier creates a straggler tail on a slow one.
 */
function groupIntoChunks(items: ScriptItem[], maxChunks: number): Chunk[] {
  type Base = { rawText: string; pauseAfterSec: number }
  const base: Base[] = []
  for (const it of items) {
    if (it.type === 'text') base.push({ rawText: it.text, pauseAfterSec: 0 })
    else if (base.length > 0) base[base.length - 1].pauseAfterSec += it.seconds
    // a leading pause with nothing to attach to is dropped — nothing to hold silence before
  }

  if (base.length <= maxChunks) {
    return base.map((b) => ({ rawText: b.rawText, silenceAfterSec: b.pauseAfterSec }))
  }

  const boundaryCount = base.length - 1
  const keep = Math.max(1, maxChunks - 1)
  const chosen = new Set(
    Array.from({ length: boundaryCount }, (_, i) => i)
      .sort((a, b) => base[b].pauseAfterSec - base[a].pauseAfterSec || a - b)
      .slice(0, keep),
  )

  const chunks: Chunk[] = []
  let parts: string[] = []
  for (let i = 0; i < base.length; i++) {
    parts.push(base[i].rawText)
    const isLast = i === base.length - 1
    if (isLast) {
      chunks.push({ rawText: parts.join('\n\n'), silenceAfterSec: 0 })
      break
    }
    const gap = base[i].pauseAfterSec
    if (chosen.has(i)) {
      chunks.push({ rawText: parts.join('\n\n'), silenceAfterSec: gap })
      parts = []
    } else if (gap > 0) {
      parts.push(`<pause seconds="${gap}"/>`) // re-inlined for the intra-chunk transform
    }
  }
  return chunks
}

// ── 3. Per-tier intra-chunk transform: what a pause becomes INSIDE a chunk ─

interface VoiceTier {
  id: string
  /** Does this tier honor a literal pause marker natively, or does it need
   *  translation to inline cue tags ("[short pause]" style)? */
  honorsNativePauseMarker: boolean
  /** Stability/consistency setting. Raise this for any tier that synthesizes
   *  chunks INDEPENDENTLY with no cross-chunk context — low stability lets
   *  the voice "wander" (accent, persona) chunk to chunk (RECIPE.md: accent
   *  drift failure mode). */
  stability: number
}

function intraChunkText(rawText: string, tier: VoiceTier): string {
  if (tier.honorsNativePauseMarker) {
    // Tier reads the marker natively — leave it alone.
    return rawText
  }
  // Tier needs translation to inline cue tags it actually understands.
  return rawText.replace(/<pause seconds="(\d+(?:\.\d+)?)"\s*\/?>/gi, (_, sec) =>
    parseFloat(sec) <= 1.5 ? ' [short pause] ' : ' [long pause] ',
  )
}

// ── 4. Synthesize chunks, bounded by the provider's concurrency ceiling ───

/** Run `fn` over `items` in waves of at most `concurrency`, preserving order. */
async function runInWaves<T, R>(
  items: T[],
  concurrency: number,
  fn: (item: T) => Promise<R>,
): Promise<R[]> {
  const out: R[] = []
  for (let i = 0; i < items.length; i += concurrency) {
    const wave = items.slice(i, i + concurrency)
    out.push(...(await Promise.all(wave.map(fn))))
  }
  return out
}

/**
 * Synthesize one chunk with an OWN, strongly-referenced timeout chained to a
 * caller abort signal. RECIPE.md step 5: a composite/derived AbortSignal with
 * nothing holding the timeout strongly can silently never fire in some
 * runtimes, hanging the request and starving a shared concurrency slot.
 */
async function synthesizeChunk(
  text: string,
  tier: VoiceTier,
  opts: { signal?: AbortSignal; timeoutMs: number; ttsCall: (text: string, tier: VoiceTier, signal: AbortSignal) => Promise<Buffer> },
): Promise<Buffer> {
  const ac = new AbortController()
  const onExternalAbort = () => ac.abort()
  if (opts.signal) {
    if (opts.signal.aborted) ac.abort()
    else opts.signal.addEventListener('abort', onExternalAbort, { once: true })
  }
  const timer = setTimeout(() => ac.abort(), opts.timeoutMs) // explicit, strongly held
  try {
    return await opts.ttsCall(text, tier, ac.signal)
  } finally {
    clearTimeout(timer) // always clear — this is what makes the timeout GC-safe
    if (opts.signal) opts.signal.removeEventListener('abort', onExternalAbort)
  }
}

// ── 5. Reinject real silence at chosen boundaries, concat, normalize ──────

/**
 * Interleave chunk audio with reinjected silence, then concatenate by
 * RE-ENCODING (never raw byte-copy — RECIPE.md step 6) and fold loudness
 * normalization into that same encoding pass (step 7). `concatAndMaster`
 * is a stand-in for your own audio toolchain (e.g. an ffmpeg invocation
 * with a concat demuxer + a loudnorm filter in one pass).
 */
async function assembleFinalAudio(
  chunks: Chunk[],
  chunkAudio: Buffer[],
  deps: {
    makeSilence: (seconds: number) => Promise<Buffer>
    concatAndMaster: (buffers: Buffer[], loudnessTargetLufs: number, truePeakDb: number) => Promise<Buffer>
  },
  loudnessTargetLufs = -16,
  truePeakDb = -1.5,
): Promise<Buffer> {
  const interleaved: Buffer[] = []
  for (let i = 0; i < chunkAudio.length; i++) {
    interleaved.push(chunkAudio[i])
    const gap = chunks[i].silenceAfterSec
    if (gap > 0 && i < chunkAudio.length - 1) {
      // Stretch slightly and clamp — a proven pacing policy, not a magic number
      // you need to invent from scratch. See RECEIPTS-internal.md for the
      // exact multiplier and clamp used at origin.
      interleaved.push(await deps.makeSilence(gap))
    }
  }
  return deps.concatAndMaster(interleaved, loudnessTargetLufs, truePeakDb)
}

export { parseScript, groupIntoChunks, intraChunkText, runInWaves, synthesizeChunk, assembleFinalAudio }
export type { ScriptItem, Chunk, VoiceTier }
reference/failure-ladder.md2.1 KB
# Reference: the failure ladder (RECIPE.md step 8)

A shape to adapt, not a fixed list — name every real failure point in YOUR
pipeline, give each a wall-clock detection budget, and give each a fallback
action. The bottom rung is always a plain path that already works.

| Point | Detection | Action |
|---|---|---|
| Upstream text generation stalls before any output | error/no-first-token event | abort this attempt, fall back to batch |
| Upstream text generation emits nothing within budget | wall-clock timer | abort, fall back to batch |
| TTS connection fails to open within its connect budget | connect timeout | switch to the provider's plain HTTP/single-shot mode, if one exists |
| TTS returns no audio within budget from stream start | wall-clock timer | abort, fall back to batch |
| Mid-stream error from either upstream | error event | abort and fall back — UNLESS enough audio has already reached the client that finishing cleanly (a replayable partial) beats restarting from zero |
| Client-side stalled read | player/transport error | client-side retry against the plain batch endpoint |
| A synthesis request hangs past its own timeout | the per-chunk timeout in `reference/chunk-and-pace.ts` | abort that specific chunk, free its concurrency slot immediately — do NOT let the whole generation wait on it |
| Provider returns a rate-limit response | HTTP 429 / equivalent | retry with backoff bounded by the same request's timeout, do NOT immediately dump the whole generation to the fallback tier |

## The rule this table encodes

Every row needs BOTH a detection (how do you know this failed, and how
quickly) and an action (what happens instead). A row with detection but no
tested action is a failure mode waiting to become an incident. A row with
neither is invisible until a user reports it.

The last rung — plain batch fallback — must be a path that is ALREADY
proven and already running for some traffic (e.g. the pre-streaming
implementation you had before this recipe). Never write new "just in case"
fallback code that itself has never been exercised in production; an
untested fallback is not a fallback, it's a second failure mode.

Where this comes from

Abstracted from a production voice pipeline that turns AI-authored narration into streamed audio for a live, latency-sensitive listening experience. It shares patterns and shapes, with no source code and no provider or model names.

On this page