Free resources

// Toolkit · 10 repos

Build a full Claude Code agent system

You saw the carousel. Here's the full map: 10 GitHub repos, one per step. The graph coordinates the fleet. The loop makes each node trustworthy. Every node in the graph is one agent running a loop. Bookmark this one.

NG
Nishant GuptaAI Systems Architect
Sep 9, 2026
The 10-repo Claude Code agent stack

// THE MAP, IN ONE LOOK

Two layers, one screen. The graph coordinates the fleet: who runs, and when. The loop makes each node trustworthy: whether you can believe what it sends back.

The graph · coordinate the fleet

bernsteinschedules the fleet
agent-worktreeisolates each agent
wshobson/agentsgives each node a role
insane-researchthe whole graph, already one plugin

Now zoom into any one of those nodes. Every node in the graph is one agent running a loop.

The loop · make one node trustworthy

every graph node runs this loop trust, not just output beadsmemory wakuthe loop core serenacontext superpowersskills review-loopthe gate workshopproof

The graph decides who runs and when. The loop decides whether you can trust what comes back. Build a graph out of loops you can’t trust, and you just ship bugs faster.

16,000 tokens of your context, gone before you type a word, on built-in tool descriptions you can’t see or edit. serena (L3) below is the repo that takes them back.

// THE GRAPH · COORDINATE THE FLEET

G1

bernstein

Schedules your fleet as a DAG in plain Python, zero model in the coordination loop. Every node hits the real Claude CLI in its own git worktree, behind a lint/type/test gate, and merges only if it's green. Trap: heavy platform, and the DAG is hand-authored, not inferred for you.

726 ★ · Apache-2.0

GitHub
G2

agent-worktree

A git worktree per agent, so five agents in one repo stop trampling each other's files. Dry-run merges first and rolls back on conflict. Trap: it's a primitive, not a scheduler. It doesn't decide how many agents to spawn.

267 ★ · MIT

GitHub
G3

wshobson/agents

Gives each node a role. A real Claude Code plugin marketplace: 203 specialist subagents across 94 plugins, tiered Opus for architecture, Haiku for fast ops. Trap: install only the roles you'll use, or you burn the context you just saved.

38,185 ★ · MIT

GitHub
G4

insane-research

The whole pattern, already shipped as one plugin. A 7-phase research graph that fans out sub-agents, then lets a deterministic code gate score every claim. Trap: skip the validate step and synthesis has nothing allowed to cite.

108 ★ · MIT

GitHub

// THE LOOP · MAKE ONE NODE TRUSTWORTHY

L1

beads

Memory that survives resets. Replaces the markdown TODO with a real dependency graph over versioned SQL. It returns only unblocked tasks, and insights survive sessions. Trap: don't put its folder on iCloud or Dropbox, cloud sync corrupts the database.

25,603 ★ · MIT

GitHub
L2

waku-agent

The loop core, small enough to read. The whole agent loop is one file, about 95 lines. No done flag: it ends when the model stops asking for tools. Trap: it swallows tool errors into normal output, so a wrong tool looks correct. That's why the gate exists.

440 ★ · MIT

GitHub
L3

serena

Context, done right. Claude Code's built-in tool descriptions eat about 16,000 un-editable tokens and bias it toward reading whole files. serena forces symbol-level retrieval: the one function, not the 2,000-line file. Trap: it verifies nothing, it only hands the model better tools.

26,813 ★ · MIT

GitHub
L4

superpowers

Skills plus discipline. A full methodology, not a pack. Its TDD skill enforces one rule: no production code without a failing test first. Trap: it's persuasion, not a syscall. A model can fake a test and claim it saw red. Real enforcement is the gate.

260,116 ★ · MIT

GitHub
L5

claude-review-loop

The gate. A Stop hook that blocks the agent from quitting until a second model reviews the work, up to four reviewers in parallel. It fails open by design. Trap: no license in the repo, and it only checks a review file exists, the agent can still skip findings.

706 ★ · no license

GitHub
L6

workshop

Proof. Captures a real run, re-drives that exact trace against your edited code, and diffs the tool calls until the spans go green. Trap: replay is only safe if you extract a clean entrypoint, or it can hit your real database and billing.

937 ★ · MIT

GitHub

// THE ONE TEST

✓

Can your system take “done” back?

bernstein won't merge a node that fails its gate. beads flips a finished task back to not-ready. The review hook un-finishes a finished session. workshop fails a green trace. A system that can only promote is a burndown chart with extra steps. Draw the graph, but build it out of loops first.

One AI automation idea. Every week.

Practical systems and breakdowns, straight to your inbox. Free. Unsubscribe anytime.

You're in. The next one lands in your inbox.