Agent Recipesby Naturate

Run a fleet of AI agents on real work without babysitting them

Several agents at once, on a schedule, while you're not watching, and it keeps going wrong: they clobber each other, run on your laptop with your credentials, stop to ask permission for everything, or fail silently for days. Here's the pattern that keeps a fleet safe and honest: isolated ephemeral workers, output-gated not permission-gated, fan-out caps, and a dumb watchdog that pages on silence.

In productionEvidence from one production system.
System
agent worker fleet
Origin
Abstracted from a production agent fleet.

Try this recipe in your own project

Read the files below first. The download for the whole kit asks for your email.

Get the kit (.zip)
  • SKILL.md
    3.6 KB
  • reference/heartbeat-watchdog.sh
    1.8 KB

Start in a test project with synthetic data. A skill file provides instructions; it does not grant permissions, start an autonomous process or establish compatibility with every agent host. Check this guide’s evidence and applicable file-specific licence.

The problem

You want AI agents doing real work, several at once, on a schedule, while you're not watching. But it keeps going wrong: two agents edit the same repo and clobber each other; they run on your laptop with your credentials; they stop to ask permission for every keystroke so nothing is actually unattended; and the worst one. They fail silently, and you discover days later that the whole fleet has been doing nothing while a dashboard cheerfully said "in progress."

What you get

A pattern for running a fleet of agent workers that stays safe and honest:

  • Workers can't clobber each other or your machine: each runs isolated and ephemeral, never on your laptop, never with standing credentials.
  • They actually run unattended: because you gate their output, not every action, only a few genuinely dangerous things stop for you.
  • Silence pages you: a dead-man's switch treats "the work didn't happen" as a first-class failure, so a multi-day blackout cannot go unnoticed.
  • Runaway fan-out is capped: depth, breadth, and concurrency limits keep a chain of agents from spawning a swarm.

Use when / skip when

Use when you're running more than one or two agents, especially on a schedule or unattended, on work that touches a real repo or real systems.

Skip when it's a single interactive agent you're watching, you don't need fleet machinery for that, and adding it is its own failure mode.

Ingredients

  • Isolated, ephemeral compute for workers, cloud sandboxes / CI runners / fresh containers. Never your primary machine.
  • Scoped, expiring credentials: per-worker tokens with the narrowest rights (branch-push, not merge; no prod, no secrets), and a spend cap on the model key.
  • A verification gate: CI (tests + build) and PR review the work flows through.
  • A dumb watchdog: a scheduled check that pages on silence. See reference/heartbeat-watchdog.sh.

The recipe

1. Never run unattended work on your own machine. Workers run in isolated, ephemeral environments, cloud sandbox or CI runner, fresh per task, torn down after. No long-lived working tree means no clobbering and no state carryover. (Schedules that depend on a personal machine miss runs. Headless jobs in CI do not.)

2. Gate dispatch manually, ask before you fan out. One accountable human triggers each batch; there is no idle auto-scheduler spawning work. Work flows: umbrella conversation → decompose into tickets with explicit done-conditions → dispatch each to one worker, scoped so they don't collide → stitch (review every landing) → digest for the human. Coordination is by decomposition, not by agents chatting.

3. Verify output, not permissions. The default guardrail is the work being checked, tests pass, build is green, the PR diff matches intent, not a human approving each action. Reserve human gates for a tiny irreducible set (typically money, production data, and anything published outward). An already-approved task is pre-authorized; don't re-gate it. Everything else: the worker decides and ships within its lane, and you review the result.

4. Cap fan-out at every layer. A wake chain that spawns wakes needs hard ceilings or it swarms: a depth limit (no spawning past N generations), a breadth limit (a single root task can pull at most K workers across its whole tree, ever), a per-dispatch concurrency limit, and, if workers share a quota, make exhaustion degrade visibly (a logged "skipped"), never silently.

5. Build a dumb watchdog, and treat silence as failure. The most important component is the dumbest: a dead-man's switch that is dumber than what it watches no AI, no dispatch, just "did the expected work post something in the last N hours? If not, page a human." If a worker ran but produced nothing, that is as bad as not running at all, wire the alarm to silence, not just to errors. See reference/heartbeat-watchdog.sh.

6. Avoid the shared-tree hazard explicitly. Ephemeral fresh-checkout-per-task is the clean fix. Where sessions do share a tree, stage with explicit pathspecs (never add -A, which sweeps a neighbor's WIP), rebase with autostash, never force past a conflict, and make the sync fail-open (offline or locked → exit cleanly, never block a session).

7. Observe on probation windows, not calendar dates. A runtime that is designed in a day and abandoned three days later was never watched long enough for anyone to learn why it failed. Give a new runtime a real probation window and resist redesigning before the evidence is in. If it keeps needing a redesign every few days, the honest conclusion is that standing autonomy is premature.

Evidence

From an agent fleet run in production.

ClaimNumber
Fan-out capssmall fixed limits, for example depth of 4 generations, 3 workers per root task and 3 concurrent per dispatch
Watchdogchecks ~1h after each scheduled slot, 3h lookback window, pages in < 1h of silence
Compute cost~$0 incremental (fits inside included CI minutes) + a per-worker, spend-capped model key
What decided the architectureruns scheduled on a personal machine were missed; headless CI jobs were not

What might go wrong

Early versions of a fleet tend to fail the same way in different costumes. This section is the real value of the recipe.

  • Control-plane accretion (the underlying disease). Every failure was answered by adding a component that reports state instead of one that does the work. The fix is never another dashboard; it's making the work actually happen and watching for its absence.
  • State theater. The board said "in progress" while the runs were failing silently. A status that isn't derived from real output is a lie waiting to happen. Bind status to artifacts, not to an agent's self-report.
  • Approval choreography. Re-gating already-approved work left a pile of tasks "approved and idle" for days. If approval isn't authorization, nothing ships.
  • Silence with no watcher. The designed failure signal was silence, and nobody was watching it, because the human was the monitoring system. Result: a multi-day blackout that nobody notices. This is why step 5 exists.
  • Laptop dependency. A schedule that quietly depends on your machine being awake isn't a schedule. If it can't run headless, it can't run unattended.
  • Redesigning before observing. Moving too fast to ever observe the failure is itself the failure. See step 7.

Keeping it healthy

  • The dumb watchdog comes first: a dead-man's switch on the fleet's output, dumber than the fleet, pages on silence.
  • Probation windows on every runtime change; evidence before redesign.
  • A one-page operating model. If keeping the fleet alive needs more than a page of rules, the design is probably wrong.

The skill

The skill and its reference code, in full. This is what an agent follows to build the pattern in your codebase. Copy any file here, or get every guide as one kit from naturate.io/recipes.

SKILL.md3.6 KB
---
name: agent-worker-fleet
description: Run a fleet of AI agents on real work without babysitting them — isolated ephemeral workers, output-gated (verify don't permission), fan-out caps, and a dumb watchdog that pages on silence. Use when running more than one agent, scheduled or unattended, on real repos/systems. Heavy on hard-won failure modes.
---

# Set up a safe agent worker fleet

You are helping a developer run multiple AI-agent workers safely — scheduled or
unattended, on real work. Adapt to their compute (CI, cloud sandboxes, containers)
and their repo host. Lead with the failure modes: most of this recipe's value is
in what NOT to do. Do not oversell — the pattern is proven by three failures; a
polished finished runtime is not.

## Detect their setup

Where workers would run (their CI / cloud / would they be tempted to run on the dev
machine — steer them off it), their repo host + how they'd scope worker tokens,
whether they have any scheduled/unattended ambition or just parallel interactive
agents (if the latter, tell them they don't need fleet machinery yet).

## Build this

1. **Isolated ephemeral workers.** Each worker runs in a fresh cloud sandbox / CI
   runner, fresh checkout per task, torn down after. Never the dev's machine, never
   a long-lived working tree. This is non-negotiable — schedules that depend on a personal machine miss runs; headless CI jobs do not.

2. **Scoped, expiring credentials.** Per-worker tokens with the narrowest rights
   (branch-push, not merge; no prod; no secrets), and a spend cap on the model key
   so a runaway worker can't run away with cost.

3. **Manual dispatch, one accountable human.** No idle auto-scheduler. Flow: talk →
   decompose into tickets with done-conditions → dispatch each to ONE worker,
   scoped not to collide → review every landing → digest.

4. **Verify output, not permissions.** CI (tests+build) + PR review is the
   guardrail. Reserve human gates for a tiny irreducible set (money, prod data,
   outward publishes). Don't re-gate already-approved work — approval is
   authorization.

5. **Fan-out caps.** Depth ceiling (no spawning past N generations), breadth
   ceiling (≤ K workers per root task, ever), per-dispatch concurrency limit, and
   visible degradation on shared-quota exhaustion (never silent).

6. **A dumb watchdog on silence.** Dumber than what it watches — no AI, no
   dispatch. Checks "did the expected work post something recently?" and pages a
   human if not. Silence is a top-level failure, equal to an error. See
   `reference/heartbeat-watchdog.sh`.

7. **Shared-tree safety** (if any tree is shared): explicit-pathspec staging (never
   `add -A`), rebase with autostash, never force past a conflict, fail-open sync.

## Guardrails to enforce

- Nothing unattended on the human's machine.
- Every status must derive from real artifacts, never an agent's self-report
  (a status board can say "in progress" while nothing is running).
- Answer a failure by making the work happen + watching for its absence — never by
  adding another component that only reports state (control-plane accretion).
- New runtime → probation window, evidence before redesign.

## Verify before done

Confirm a worker runs with no access to prod/secrets and can only push a branch.
Kill a scheduled run and confirm the watchdog pages within its window. Confirm two
workers on different tasks don't touch the same paths. Confirm exhausting the quota
logs a visible "skipped," not silence.

## Read alongside

`RECIPE.md` (the why + the three-deaths failure history) and
`reference/heartbeat-watchdog.sh`.
reference/heartbeat-watchdog.sh1.8 KB
#!/usr/bin/env bash
# Reference: the dumb watchdog (step 5). A dead-man's switch on the fleet's OUTPUT.
# Deliberately dumber than what it watches: no AI, no dispatch — just "did the
# expected work post something recently? if not, page a human." Silence is the
# only trigger. Run it on a schedule ~1h after each slot your fleet is meant to act.
#
# This example watches a GitHub issue used as the fleet's log/heartbeat. Adapt the
# "did work happen?" probe to wherever your workers leave evidence (a log channel,
# a metrics ping, a status file, a DB row).
set -euo pipefail

REPO="${REPO:?set REPO=owner/name}"
HEARTBEAT_ISSUE="${HEARTBEAT_ISSUE:?issue number the fleet posts its log to}"
LOOKBACK_HOURS="${LOOKBACK_HOURS:-3}"      # fresh = posted within this window
PAGE_HANDLE="${PAGE_HANDLE:-@your-oncall}" # who gets mentioned on silence

# Cutoff = now - LOOKBACK_HOURS, in ISO-8601 UTC.
cutoff="$(date -u -v-"${LOOKBACK_HOURS}"H +%Y-%m-%dT%H:%M:%SZ 2>/dev/null \
  || date -u -d "${LOOKBACK_HOURS} hours ago" +%Y-%m-%dT%H:%M:%SZ)"

# The probe: any fleet activity on the heartbeat since the cutoff?
recent="$(gh api "repos/${REPO}/issues/${HEARTBEAT_ISSUE}/comments" \
  --jq "[.[] | select(.created_at > \"${cutoff}\")] | length")"

if [ "${recent}" -gt 0 ]; then
  echo "OK: ${recent} fleet post(s) since ${cutoff} — fleet is alive."
  exit 0
fi

# SILENCE. That is the failure. Page, and exit non-zero so the scheduler
# (CI, cron) also surfaces it through its own failure channel (email/alert).
echo "SILENCE: no fleet activity since ${cutoff}. Paging."
gh api "repos/${REPO}/issues/${HEARTBEAT_ISSUE}/comments" \
  -f body="${PAGE_HANDLE} FLEET SILENT — no activity in ${LOOKBACK_HOURS}h. The work did not happen (or ran and posted nothing). Check the workers." >/dev/null
exit 1

Where this comes from

Abstracted from a production agent fleet. It shares patterns and lessons, with no source code.

On this page