Drawing Mode (Press 'C' to clear)
NerdyNight
NerdyNight

Loop Engineering

Infrastructure that lets coding agents work autonomously at scale — while engineers stay in control of what ships

Varunkumar Nagarajan

Agenda

"Build the loop. Automate the toil. Stay the engineer who decides what ships."
01 Building the Loop Six questions, six pieces, then the definition
02 Loops in Action Use cases, a real session
03 Lessons & Close What worked, what's next, discussion
Chapter 1
Building the Loop
One need at a time — the definition comes last

How many sessions do you run in parallel, right now?

One tab? Three? A window for every ticket?

Today: One, Full Attention

A ticket lands. You open Claude Code and start a conversation:

> Here's the ticket. Client sees a null FX rate on book 4567.
Reproduce it. Find the root cause. Write a fix.

You steer it step by step: pointing it at the right repo, course-correcting when it drifts, approving tool calls, testing changes, reviewing the MR.

  • You're the scheduler that decides when the next one starts
  • You're the safety net on every tool call
  • You're the budget tracker
  • You're the notification system

Four jobs, one you. That's why it's one session at a time.

Isolation

For a second session to run without stepping on the first, each one needs its own workspace, its own process, its own budget:

common/bin/heartbeat.py
def build_launch_cmd(item, worktree, runner_model, plugin_dir, entry_skill, session_id=None, mcp_config_path=None) -> str: # pins the session to a known uuid so a stuck # runner can later be resumed interactively prompt = f"Run {entry_skill} for {item.id}" ... return ( f"cd {worktree} && " f"claude -p --model {runner_model} " f"--permission-mode auto {session_flag}" f"{mcp_flags}{plugin_flags} " f'"{prompt}"' )

One OS process per claimed item. Its own git worktree, its own budget cap, zero inherited MCP servers. Now you can run more than one — if something starts them.

So — who starts the next one?

You do. Every time.

Scheduling & Discovery

Something has to notice work exists and decide to start it, deterministically, on a schedule:

  • Git pull the state repo
  • Guardrail preflight — self-test the safety check
  • Discover work items (no model call, no token spend)
  • Triage and filter — drop anything already claimed
  • Claim by writing a file and pushing to git
  • Set up an isolated worktree and launch a session
  • Post a Slack summary to the loop channel

The model is never invoked during the heartbeat.

Zero AI Spend: the Control Plane

common/bin/heartbeat.py
def main(loop: str, deps: HeartbeatDeps) -> dict: """Run one heartbeat tick for the given loop. Steps (ยง5.2): 1. Plugin-version gate (before touching anything). 2. git pull --rebase. 3. Preflight (fail-closed). 4. Discover work items. 5. Filter to new items (no claim, no session). 6. Cap-gated claim loop. 7. For each won item: setup, write QUEUED session, launch, mirror. 8. Refresh heartbeat on owned active claims. 9. Post LOOP_SUMMARY. Always runs. """

The claiming, budget math, and safety self-test are all plain Python. No LLM interaction anywhere in this file.

Why Not Just Use Native Scheduling?

Both Codex and Claude Code now ship their own scheduling. I keep the trigger outside the agent on purpose:

Codex Claude Code (cron / routines) This approach
Where the trigger lives ChatGPT web/desktop — not the CLI Inside the agent runtime itself A plain Python heartbeat, no agent involved
Who decides "is there work?" Hosted Codex Cloud, opaque The model, via a tool call, every tick A deterministic script — zero model calls
Token spend per idle tick Unknown, hosted > 0 — the model reasons every wakeup $0 — nothing to spend on
Safety rules Server-side, undocumented Per-session hooks One fail-closed guardrail, shipped in the plugin

The scheduler doesn't need to be smart. It needs to be free.

Now a few of these are running unattended. What if one tries git push --force on main? Reads your SSH key?

Would you even know before it happened?

Guardrails

A permission boundary that doesn't depend on the model behaving — fail closed, not fail open:

common/hooks/guardrail.py
_DEFAULT_DENY_BASH = [ r"(^|\s)aws(\s|$)", # no aws CLI r"git\s+push\b.*(\s|:)main\b", # no push to main r"git\s+push\b.*(-f\b|--force\b)", # no force-push ] def main() -> None: try: ... except Exception: # any unhandled error becomes a deny _deny("guardrail error: failing closed")
  • Deny list ships inside the plugin — every developer inherits the same protection
  • An unhandled exception in the hook itself becomes a deny, not an allow
  • The heartbeat self-tests the guardrail every tick before claiming work

What if one just… keeps going? Burning tokens all night?

Who's watching the meter?

A Budget That Isn't a Suggestion

common/hooks/budget.py
def _handle_pre_tool_use(loop_config, session, transcript_path, tool_name): # deny every tool except Slack once halt_pct hits spend = compute_spend(transcript_path, price_table) pct = spend / cap * 100.0 halt_pct = budget_cfg.get("halt_pct", 100) if pct >= halt_pct: if tool_name == slack_tool: return # let it post BUDGET_HALTED _emit_deny(pct)

Warn at 70% and 80%. Halt at 100%, blocking every tool except the Slack notification. You decide whether to extend it.

You're not staring at any of these terminals anymore. So how do you know what happened?

Do you remember to check back?

Shared State & Notification

Something durable that survives you not watching, and a place status actually shows up:

GIT REPO

  • Append-only audit log, one file per item
  • Every phase transition, every event, timestamped

SLACK + DASHBOARD

  • One thread per work item, updated live
  • A dashboard rendered by CI from the same repo

You stop watching terminals and start reading a thread.

This works for one kind of ticket. What about something completely different — a dependency upgrade?

Do you rebuild the whole thing?

A Loop Is Just a Config File

loops/data-inconsistencies/config.yaml
loop: data-inconsistencies phases: [TRIAGE, SCENARIO, REPRO, RCA, FIX, VALIDATE, REVIEW, MR_RAISED] entry_skill: investigate-data-inconsistency enabled: true discovery: source: jira_filter params: jql: 'labels in (loop-eligible) AND assignee is EMPTY' workspace: type: worktree base: origin/main caps: { global: 5, per_user: 4 } budget: { total_usd: 30, warn_pct: [70, 80], halt_pct: 100 }

Nothing in the scheduler, isolation, guardrail, budget, or shared-state layer changes when you add one of these.

Six Needs, Two Planes

CONTROL PLANE

  • Scheduling & discovery
  • Distributed claiming (git-optimistic lock)
  • Deterministic Python — zero AI spend
→

EXECUTION PLANE

  • Isolation — one worktree, one process
  • Guardrails — fail closed
  • A budget that halts

SHARED STATE, AND A CONFIG FOR EACH KIND OF WORK

  • Audit log, Slack thread, dashboard
  • Every loop: the same infrastructure, a different YAML

Six questions. Six pieces. All of it running right now.

You Just Built One

A scheduled, long-running system that picks up a unit of work, carries it start to finish, and hands a finished artifact to a human for review.

It usually doesn't merge or ship. It produces a result for an engineer to judge.

prompting agents by hand → building the system that prompts them

This follows the "Loop Engineering" approach described by Addy Osmani. The central discipline: automate the toil, not the judgment.

Chapter 4
Loops in Action
Use cases and a real session

The Infrastructure Is Loop-Agnostic

A loop is a YAML config and an entry skill. Nothing else changes.

Data Inconsistency Resolution

Nine phases: scenario, repro, RCA, fix, 20x validation, review, MR. Hard gate: must reproduce before it can fix.

Alerting System Triage

Investigates the underlying cause and surfaces a triage summary. Touches no code. Three phases.

Dependency Upgrade

Scans manifests for outdated/vulnerable deps, creates upgrade branches, runs tests, raises MRs.

Security Findings

Triages severity, identifies affected code paths, drafts remediation patches for review.

The Heartbeat in Motion

  • Discover eligible tickets via a direct MCP client
  • Triage & filter against per-machine and global caps
  • Claim by writing one file and pushing — no central lock server, just git-optimistic locking
  • Set up worktree, mark QUEUED, launch the session
  • Post a Slack summary with the tick status

If another machine claims it first, the loser yields. No coordination overhead.

First Live Session

2026-06-29
#loop-updates
DATA-24280 — Record inconsistency. Null last_sync_timestamp, mismatched values between two replicas, dataset shard 4567.
Session completed triage, transitioned to scenario generation. Potential RCAs found before the issue failed to reproduce.
  • Claiming, worktree setup, state tracking, notifications, guardrails — all worked as intended
  • The debugging skill itself needs refinement
Chapter 5
Lessons & Close
What worked, what's next, and over to you

Automate the toil.
Not the judgment.

"Verification is still on you. A loop running unattended is also a loop making mistakes unattended."
— Addy Osmani

The Road Ahead

  • Complete the first data-inconsistency session end-to-end through all nine phases
  • Validate the multi-machine scenario beyond a single pair of dev environments
  • Stand up a second loop — test coverage — to prove the infrastructure generalizes without changes

Currently: proof of concept, unit-tested.

Beyond One Loop: Graph Engineering

Loops connected to loops, with structure in the connections.

One loop drifts eventually — metrics get gamed, targets go stale, nobody watches the watcher. The fix isn't a smarter loop, it's topology:

  • Watcher loops that monitor what other loops do
  • Oversight loops that question whether targets still hold
  • Conflict-resolution loops when objectives collide
  • Audit loops whose only job is checking that the numbers still touch reality

None of it works without anchors — immovable ground truth (real revenue, a real reproduction, a real diff) that resists getting gamed. Today we have one loop. The graph is what comes after the second and third.

Varun
We've spent decades automating deterministic work. What changes with loops is that the worker inside them is no longer a dumb script — it's an intelligent agent that can read a ticket, investigate, and improvise.

That's exactly why the infrastructure around it is deliberately boring, predictable, and fail-safe. The more capable the agent, the more that separation matters.

Build the loop. Automate the toil.
Stay the engineer who decides what ships.

Thank You!

Varunkumar Nagarajan

Let's talk — your use cases, your loops.