Madav UniversityA creator's blog, shared forward

From Idea to Impact

I did not begin with AI.
I began with the problem I wanted to solve.

I was not waiting for mastery. Building was how I began to understand.

AI in Plain Words

This is how I came to understand AI,
drawn as one picture.

From the chat box to the whole picture, one plain word at a time.

Git in Plain Words

What Claude just did
to your code.

The story of one small change, from start to finish.

This is the honest path from a first repository I barely understood to a product shaped by hundreds of experiments. I am sharing it so your first steps can be clearer than mine.

Before Madav, GitHub repositories, VS Code and the terminal were unfamiliar territory. I opened open-source projects, followed the instructions without knowing every word yet, watched what changed, and asked a better question each time something failed. Prompts, context, memory, retrieval, tools and agents first arrived as separate problems. As the experiments began to depend on one another, the playground became Madav. AI stopped looking like a room reserved for developers. Domain experience gave me direction: I knew which problems mattered and what a useful result should feel like. AI tools supplied leverage while technical fluency grew through the work. Neither replaces the other.

How I started learning AI

The model, Claude, GPT, is the brain. The harness is the body you build around it. Its senses, its memory, its hands, its judgment. Same brain, no body, useless. Same brain, great body, unstoppable. So the real skill isn't picking a brain, it's building the body. It's called the harness. It's built from your decisions, your corrections, your understanding of business.

Level one is the chatbot. You type, it answers.

It's the doorway, not the house.

The same brain
And here's the thing about it: it's the same for everyone. You're all talking to the exact same brain.
Forgets you
Brilliant in the moment, and it forgets you the second you close the tab.
The myth
Most people think prompt engineering is what AI is.Flip it ↻AI is more than a chat screen for daily operations. Start with Claude.Flip back ↻

My recommendation

What I did, as three steps you can take.

01

Start with Claude

Claude is more than just a chat screen. Complete the course for a good understanding of Claude.

03

Build with Claude

Build Madav Coffee Cup's four helpers with Claude, one file at a time.

Why this page exists

Nobody explained this to me in one place, so I put it in one place. Every part below is written the way I wish someone had written it for me: no jargon, nothing assumed, and no need to learn it all at once. Read the day at the shop, see the same day drawn as one picture, then take any part by name.

A Day at Madav Coffee Shop

Everything about AI, told in one shift.

Madav Coffee Shop hired a helper this week. It can write, count, plan, read a delivery note and hold a conversation. It has never worked in a shop before, and it will forget everything by tomorrow morning. That is not a fault. It is the most important thing to understand about it.

This is a day in its life. By the end, you will know how AI actually works, and you will never have read a definition.

A cutaway drawing of Madav Coffee Shop with every object labelled: the gate, the desk, the binders, the button, the regulars card box, the supplier hatch, the sockets, the library, the SEND lever and the bell
The shop, and everything in it. Every object here is one idea, and you will meet them all before closing time.
  1. The Madav robot reading a HOUSE RULES board in the early morning, one hand on a button labelled slash open-shop, the door sign turned to Open

    7:45

    The shop opens

    The helper arrives before the light does. The first thing it does is look up at the board on the wall, the one bolted there by the owner: be kind, never share a customer’s details, refunds up to £10, if unsure ask the manager. It reads those four lines the way a new member of staff reads them on a first morning, and it will read them again before every single thing it does today.

    Then it presses one button on the counter, and the shop wakes up: lights, machine, specials board, the sign on the door turned to OPEN. One press, the whole routine.

    This is System instructions and Commands.

  2. The Madav robot at a small desk buried in order slips, older slips sliding off the edge onto the floor where a note reads Allergic to nuts, while it offers an almond croissant

    9:10

    The desk fills up

    By mid-morning there is a small wooden desk behind the counter, and everything the helper is working from is on it: the house rules, the menu, the last few messages from the customer it is serving. It can only work with what is on that desk. It has no filing cabinet in its head.

    The desk is small. New slips arrive from the right, and as they land, old ones slide off the left edge onto the floor. One of the papers on the floor says ALLERGIC TO NUTS.

    Sam wrote that at the start of the conversation, an hour ago. The helper is now smiling, holding out an almond croissant, and it has no idea. It did not forget the way people forget. The note simply is not on the desk any more.

    This is Context window and Hallucination.

  3. A brass gate down across the serving hatch signed Before serve: allergy check, stopping a tray with an almond croissant while a NUT ALERT card pops up

    9:11

    Someone catches it

    A brass gate drops across the serving hatch. It drops every time. Not because the helper remembered, not because it was having a careful day, but because the owner bolted a rule to the hatch itself: before serving food, check the allergy card. A little card springs up. NUT ALERT. The tray stops.

    Sam gets a plain muffin and never learns how close that was. That is the difference between telling someone to be careful and making it impossible to be careless.

    This is Hooks and Guardrails.

  4. The Madav robot laying the same desk with a mat of four named slots and pinning a NO NUTS card into the second, beside a REGULARS card box and a shelf of labelled binders

    11:30

    The desk gets laid properly

    The owner does not buy a cleverer helper. She buys a placemat. Now the desk has marked slots, filled in the same order every time: the rules, this customer, the recent messages, and the new message last. Sam’s card — oat milk, no nuts — is pinned so it can never slide off again. Everything else stays on a trolley marked not needed today.

    Beside the desk: a card box of regulars, so the helper starts each morning knowing Sam again. And a shelf of binders — latte art, refunds, catering — which it takes down one at a time, only when the job calls for it.

    Same helper. Better morning. All that changed was what it was given.

    This is Context builder, Memory and Skills.

  5. The Madav robot between a supplier hatch, a panel of five identical sockets and a pegboard of labelled instruments

    14:00

    Reaching the world

    A helper that can only talk is a helper that can only talk. This one is given three ways to reach past the shop door.

    A hatch in the back wall, where the bean supplier takes a filled-in form and hands back ten bags. A panel of identical sockets, so the calendar, the bank and the email all plug in the same way. And a pegboard of labelled instruments — check stock, book a table, send email — each tagged with exactly what it needs.

    It takes down one instrument at a time. It never reaches for the whole board at once.

    This is API, MCP and Tools.

  6. A conveyor of orders under four numbered station signs on one side, and on the other the Madav robot holding a goal card inside a loop marked Look, Decide, Act, Check

    15:20

    The line and the loop

    There are two ways to get work done here, and the shop uses both.

    On one side of the room, a conveyor. Every online order travels the same four stations in the same order: read, check stock, make, send receipt. Hundreds a day, never a surprise. At the end sits a small tray labelled doesn’t fit — ask a human, for the order that says “deliver to the hospital ward and leave it at reception”.

    On the other side, no conveyor at all. The helper is holding a card that says 40 cups by Saturday, and nothing else. It works out its own steps, and then it stops, because the price is over its limit, and rings the bell for the manager.

    A route somebody drew, or a goal somebody set. Knowing which one a job needs is most of the skill.

    This is Workflow and Agent.

  7. Evening in the back office: the Madav robot with a row of numbered tasting cups, a printed step-recorder tape, a meter reading This month 340 of 500, and a clipboard offered to a human hand

    21:00

    Closing time

    The shop is dark and the real work of running the thing begins.

    Five numbered cups on the bench: the same test the helper is given after every change, so nobody finds out from a customer that it got worse. A long paper tape showing every step it took today — including the moment the supplier call failed and it answered anyway. A meter on the wall reading £340 of £500. And one clipboard held out to a human hand with a pen: please approve.

    The helper did hundreds of things today on its own. It is asking about one.

    This is Evaluation, Observability, Cost and Human in the loop.

What today was really about

Nobody in this story was clever or stupid. The helper that offered Sam an almond croissant and the helper that got the morning right were the same helper, on the same day, with the same abilities. The only difference was what was on the desk, and who had decided what belonged there.

That is the whole thing. Every idea below — every clever word you have heard about AI — is a detail of that desk: what goes on it, where it comes from, who checks it, and what happens before anything leaves the building.

You have just met all of them. Now meet them by name.

The same day, drawn as one picture

Every box below is one of the objects you just met. Follow one request from beginning to end, at whatever pace suits you, and press any box to read its part in full.

What I first saw

The chat box is only the front door.

When I began, this simple exchange looked like one magical box. Building Madav taught me to see the understandable parts behind it. You do not need to learn them all at once. Just follow one request from beginning to end.

Six words for one requestThe short version
Step 01Steps 02 - 06Steps 07 - 12Steps 13 - 17Steps 18 - 20Steps 21 - 22
Application or platform preparationPreparation layer
Model runtime - cloud or localProvider or self-hosted
Orchestration, action and outcomeApplication or platform layer
Optional reuseOver the window: stop, reject, reduce or compactIts dials: how much variety, how much thinking. What it learned stops at a date.Or it failsAnswer, ask or stop proposalRead back next time as Selected memory (01)Every step above leaves a trail: what was sent, what came back, how long it took, what it cost.Proposed actionRepeat or retry: the result becomes context for a new call - as information, never as instructions

A book that remembers

If Git is new to you, read on from here; if you use it every day, skip Getting ready. If you know the words but not the inside, read the sentence starting Underneath in each moment.

Picture Madav Coffee Shop with one notebook and a pencil. Someone rubs out the flat white and writes a new one. It tastes wrong, and nobody remembers the old recipe. Then two people change the same recipe on the same morning, and one loses their work without knowing.

Keeping every version instead, with who changed what, when and why, is called version control. Git is a free program that does it on your own computer. GitHub is a website that keeps a copy of your project for other people and machines. Git is the program; GitHub is the place.

In this story, the recipe book behind the counter is your code. Head office is GitHub, where the second copy is kept. Two other shops stand for everyone else who works from that copy. The robot is Claude. It writes; you decide what goes in the book.

Your project folder with its whole history is a repository, or repo: the book, each saved version a page. The history lives in a hidden folder inside it called .git. Delete that, and the folder on your machine forgets its past.

Getting ready, once

Do this yourself; afterwards Claude does the Git part.

  • First, install Git. On Windows, get it from git-scm.com and keep every default; on a Mac, find Terminal, the window for commands, with Command-Space, type git --version and press Install if asked.
  • Then register at github.com: press Sign up and give an email, password, username and the code GitHub emails you.
  • Next, make an empty book on GitHub: press New repository, name it, choose Private, add nothing and create it. Copy its address, ending in .git.
  • Then make a folder, add one small file, and open a terminal there. On Windows, type powershell in the folder's address bar and press Enter; on a Mac, type cd and a space in Terminal, drag the folder in and press Return.
  • Last, type each line below except the # notes, with your name, email and book address, and sign in when asked. On a Mac, do this first: install the sign-in helper (the .pkg file with arm64 in its name, or x64 if About This Mac says Intel), because GitHub refuses a typed password.
once, getting readyTerminal
# who you are, once per computer: Git writes it on every page
git config --global user.name "Your Name"
git config --global user.email "you@example.com"
# make this folder a book, and call its master book main
git init -b main
# the first page: every file, saved with one sentence
git add .
git commit -m "the first version, before I change anything"
# where head office is; origin is the usual nickname for it
git remote add origin https://github.com/you/your-repo.git
# send the book there; -u remembers the way for next time
git push -u origin main

One change, from idea to every shop

Customers say the flat white goes cold on the walk to the table. So you ask the robot for less milk and a hotter pour. Follow that one change through nine moments. Each moment ends with the lesson that tells it in full.

  1. The Madav robot beside a coffee shop counter holding a plainly bound copy of a recipe book, with the master leather book on the counter and a van leaving for head office carrying a third copy

    01

    Where code lives

    Your book lives in two places: your machine and head office. They are not automatically the same. So the robot first brings head office's newest pages home. That is a pull; sending your pages the other way is a push. Underneath, a pull is two steps: it fetches the new pages, then folds them into yours. Pull before you start, and push when you stop.

    This is Where code lives.

  2. The Madav robot writing a dated note in the margin of a recipe book page beside a changed recipe, an old crossed out line above it

    02

    Commit

    The robot makes the change: 100ml of milk, not 120, steamed to 68 degrees, not 65. It saves this as a new page with a sentence saying why. That is a commit, kept on your machine until a push. Underneath, each page is a snapshot of every recipe, and unchanged ones are stored once. Read the sentence before you approve: it is the only explanation that survives.

    This is Commit.

  3. The Madav robot at a side table working on its own plainly bound copy of the recipe book while the master lies untouched on the counter and the shop keeps serving

    03

    Branch

    But where was that page written? Not in the master book, which your project calls main. First, the robot took a clean copy to the side table and labelled it claude/flat-white-less-milk. The shop kept serving from main. That copy is a branch. Underneath, a branch, like main, is only a label clipped to the newest page, so it costs nothing. Give the robot a branch for every job.

    This is Branch.

  4. Two recipe pages laid side by side on the counter, the old one with a line struck through in red and the new one with the replacement written in green, the Madav robot pointing at the struck line

    04

    Diff

    Then the robot shows only what changed. In red with a minus, what left: 120ml and 65 degrees. In green with a plus, what arrived: 100ml and 68 degrees. That is a diff. Underneath, Git works it out from the old and new pages whenever you ask, and never saves it. Read the minus lines first: the risk is in what quietly left.

    This is Diff.

  5. The Madav robot holding out a single recipe page across the counter toward a person, the master book closed beside them, a small tray of taste-test cups between

    05

    Pull request

    The robot pushes its copy to head office and asks: may this go in? That is a pull request, a web page on GitHub, not part of Git; despite the name, nothing is pulled. Nothing is in main yet. Underneath, the push sent only the pages head office lacked. Checks, tests that run by themselves, may turn green: nothing they watch broke. It is not approval.

    This is Pull request.

  6. The Madav robot pasting an approved recipe page into the open master book on the counter while a van pulls away to carry the change to the other shops

    06

    Merge

    Then you press Merge pull request, Confirm merge and Delete branch. The page goes into the master book, and the side-table copy is thrown away. That is a merge. The other shops get it with their next delivery: a pull. Underneath, the merge adds a page pointing back to main and the copy, so deleting the copy's label loses nothing. Merging is publishing: merge when you can watch.

    This is Merge.

  7. Two recipe pages for the same drink laid over each other on the counter, both rewriting the same line differently, the Madav robot stepping back with its hands open rather than choosing

    07

    Merge conflict

    Say you had changed the same milk line on Tuesday too. Git will not guess which wins: it stops, keeps both and waits. That is a merge conflict. Underneath, Git writes both versions into the file, one above the other, until a person keeps one. Read both sides, then you choose; never just tell the robot to fix it.

    This is Merge conflict.

  8. The Madav robot lifting an older dated recipe page from the back of the master book and copying it onto a fresh page at the front, the bad recipe still in place behind it

    08

    Undo

    And if 68 degrees turns out wrong? The page that said 65 is still in the book. The robot writes it onto a fresh page at the front, with a note saying why. That is a revert, the safe way to undo. Underneath, a revert writes the exact opposite of the one bad change, so anything added since stays. Put it back first; find out why after.

    This is Undo.

  9. The Madav robot and a person looking together at a row of dated recipe pages pinned along the counter in order, each with a small note saying who asked and who approved

    09

    End to end

    By closing time the change has left a trail, your project's history. Its pages say what changed, why, and who pressed Merge; the pull request keeps what the checks said. Underneath, each page's name, like 3ca97a6, is worked out from it and every page before, so nobody can quietly change the past. After your next merge, read your trail on GitHub.

    This is End to end.

Tomorrow, Claude does it for you

When Claude makes this change for real, three lines go past. They read “created branch claude/flat-white-less-milk”, “committed: correct the flat white: less milk, hotter” and “opened pull request #1”. Then the Merge button on GitHub waits for you. You can read those lines now: a labelled copy, one page with its sentence, a request to put it in.

Behind them, Claude types lines like these for you; even its undo waits for your Merge.

the way Claude types itTerminal
# 01 bring the newest pages home; 03 take a labelled copy
git pull
git switch -c claude/flat-white-less-milk
# 02 save the change as one page; 04 show only what changed
git commit -am "correct the flat white: less milk, hotter" -m "it went cold on the walk"
git show
# 05 send the copy; 06 after you press Merge, bring main home
git push -u origin claude/flat-white-less-milk
git switch main
git pull
# if it goes wrong: 07 files in dispute; 08 undo the bad page
git status
git switch -c claude/undo-flat-white
git revert --no-commit 3ca97a6
git commit -m "back to the old flat white: 68 was wrong"
git push -u origin claude/undo-flat-white

The one job that stays yours

No tool can do it for you. Open the pull request's Files changed tab, read the minus lines first, then press Merge. If it still goes wrong, nothing is lost: put the old page back.

Now you know what Claude just did to your code.

AI in Plain Words

Models, tokens and inference

A model is a learned mathematical system that turns input into a prediction or generation.

The AI-native software lifecycle, told through one project

Building Spindle:
one chat window, any model.

Where this comes from

Anthropic published The AI-native SDLC playbook. For decades, writing code was the slow, expensive part, so every process was built to protect it: requirements, estimates, sign-offs, hand-offs. Now an agent writes most of the code in hours, and the slow parts are the human decisions around it: what to build, whether it is right, and when to ship.

Read the playbook on claude.com

Business case

This guide follows one project from its first idea into production. The project is Spindle: one chat window where a delivery team can use any model, instead of juggling four chat tools and four sets of API keys. You will watch an idea raised on a Tuesday afternoon become an intent, a spec, a plan, working code and a reviewed release, and then, when something breaks at 3 a.m., a new intent written up before anyone looks (in Spindle, a replayed night). At every stage you will see who makes the call, what Claude does, and the file that carries the work to the next stage.

Rest on any dotted term for its definition. Each stage has a flow you can read at a glance, and pressing a breathing file opens it in full.

The six-stage loop

The six stages, drawn as a loop.

An idea enters at the top, a production alert at the bottom, and both travel the same track. The centre of the drawing says the one rule that holds it together.

The traditional lifecycle is a straight line with hand-offs: a document goes to the next team, who write a different document. Here it is a loop, and the hand-off is a file in git. That chain of files is also the audit trail: who asked for what, what Claude produced, and who approved it.

The running example

The project we will follow.

One chat app, three real models and a pretend one, behind one proxy, and a team that builds it with Claude while Claude also runs inside it.

Spindle is a chat app. You type a message, pick a model from a dropdown, and Spindle forwards the message through its own proxy to that model. The proxy holds the API keys, so they never reach the browser, and it writes one log line per call. The team uses Claude to build it. Claude will also run inside it, as one of the models in the dropdown, reached through the OpenRouter adapter. That is what makes it a harness: Spindle is the wrapper around the models, not a model itself. On day one the dropdown offers Claude and GPT, reached through OpenRouter, and an open model reached through NVIDIA. A pretend model, the mock, is the default, so anyone can run Spindle with no keys. Sign-in is a simple dev account, clearly marked as a stand-in: Spindle is a teaching project, not a production system. Conversations are kept for thirty days.

Who does what

The cast.

Four people, one agent, and the checks that follow rules exactly. In every diagram below, people wear the accent, Claude wears its own orange spark, the checks wear a green tick, a trigger wears an amber bolt, and the lane says who is acting.

Mark

Leads the delivery team. He has the idea and describes what he wants in his own words.

Rahul

Product owner. He decides whether an idea goes ahead and signs off the spec.

Marcus

Engineer. He steers Claude Code, approves the build plan, and marks the pull request ready for review.

Linda

Tech lead and release manager. She judges risk, approves merges, and authorizes production.

Claude

Drafts the documents, writes the code and tests, reviews pull requests, diagnoses incidents. Never approves its own work.

The checks

Scripts, hooks, CI and monitoring. They follow rules exactly, never use judgment, and cannot be talked around.

Stage 01 – Plan – Ends by saving intent.md

Ideas stop waiting for someone to write them up.

Mark says what he wants, once, in his own words, and it becomes a file the next stage can act on.

The problem

Mark has been thinking for weeks. His delivery team pays for four different chat tools because each one talks to a different model. Prompts live in four places. So do the API keys. Every time someone wants to compare two models, they copy text between browser tabs.

The old process

In the old process this would need a request, wait for a backlog review, join a refinement meeting, and eventually a product analyst would turn his idea into user stories. By the time an engineer saw it, the idea had passed through three sets of hands and lost half its meaning.

With Claude

Now he opens Claude and just talks. "I want one chat window where you pick the model. Claude, GPT, Gemini, maybe a local one later. Our API keys must never be in the browser." Claude asks the questions an analyst would ask. Who exactly are the users? What does "any model" mean on day one? Do replies need to appear word by word, or all at once? What happens to conversation history? Twenty minutes later the idea is concrete.

Written down

Mark asks Claude to write it up as an using the company template. He reads it, fixes the one thing Claude got wrong (the local model is a nice-to-have, not a day-one need), and saves it into the intent/ folder of the Spindle code repository. He never touches git himself. Claude reads the repository through the GitHub connector, and a Claude Code session commits the file and opens the pull request for him, saying so in it.

Accepted

Rahul, the product owner, sees the new file, reads it, and accepts it. That acceptance, recorded as a merge, is the first gate in the loop. Nobody rewrote Mark's idea. The accepted file is what Rahul opens to start Stage 2.

How it flows

Three lanes: people, Claude, and the repository with its automatic checks. Click any box to read what happens there.

The old way

Idea, then ticket, then backlog, then a refinement meeting, then user stories written by someone else. Weeks pass. What reaches engineering is several hand-offs removed from what Mark meant, and nobody can trace who changed what.

The AI-native way

Mark brainstorms with Claude and saves intent.md in his own words within the hour. It says what is wanted, why, and under which constraints. The product owner accepts or closes it. The git history is the record of who asked for what, and when.

Early sign it is working

Time from Mark's first conversation to a saved intent.md. Read straight from git. It should fall from weeks to hours.

Later sign it is working

The share of intents Rahul accepts rather than closes, and how often an intent.md has to change after the spec is written.

Stage 02 – Design – Ends by saving spec.md

Requirements and design collapse into one session.

Company policy is applied while the spec is being written, not discovered in a review three weeks later.

The trigger

The accepted intent.md is the trigger. Rahul opens a Claude session that has the company's loaded: a security skill (how we handle secrets and logs), a UX skill (our design system), and a compliance skill (what we may store about people). He attaches the intent and gives one instruction: read this, produce a requirements and design spec for our codebase, apply our skills, and flag every concern, especially where two policies contradict each other.

The spec

Claude reads the intent and the existing repository, then writes a . It proposes the shape of Spindle: a chat screen, an API with a dev sign-in in front of it, a proxy and router that holds the keys, and one small adapter per model provider, all speaking the same internal format. It also flags three concerns:

  • The security skill says "never log message text", but the intent says "log every call". Which wins?
  • Streaming replies works differently for each provider. Supporting it for everyone on day one adds risk.
  • Storing conversations in our database may count as personal data under the compliance skill.

Concerns settled

In the old process, these would surface in a security review weeks after design was declared done. Here they surface before any engineer has seen the work. Rahul takes the first concern to the security lead (answer: log who, when, which model, and how many tokens, never the text). He decides streaming is day-one for Claude and GPT only. Compliance confirms a thirty-day retention rule. He writes each answer into the spec.

The mock

Because the chat screen is front-end work, Rahul also opens , points it at the intent, and iterates on a mock of the chat window with the model dropdown until it looks right. That mock is exported from Claude Design as a standalone page, design/chat-mock.html, with a hand-off prompt for Claude Code. A screenshot of it, design/chat-mock.png, is what the screenshot check compares against.

Accepted

Rahul saves spec.md next to intent.md. The pair now records what was asked for and what was decided. Because the proxy handles secrets, the company classes it as higher risk, so Rahul brings in Linda, the tech lead, before accepting. Together they accept the spec. That acceptance is the gate, and it starts Build.

How it flows

The concerns are the point: Claude raises them, people resolve them, and the answers are written into the spec.

The old way

Analysts turn the idea into formal requirements. Designers then read those requirements and turn them into a design. Two teams, two documents, and the security review happens after both are finished. Accountable, but slow, and meaning leaks at every hand-off.

The AI-native way

One session. Claude writes requirements and design together, constrained by the company's skills, with concerns flagged up front. Rahul reviews the spec rather than writing it, and resolves each concern with the person who owns that policy before engineering starts.

Early sign it is working

Time between the intent.md commit and the spec.md commit for the same idea. Two git timestamps, compared with the old requirements-plus-design cycle.

Later sign it is working

How often the spec has to change after building starts. Count spec.md commits dated after the first plan.md. It should be close to zero.

Stage 03 – Build – Ends by saving plan.md, then the code and its tests

Nothing is implemented without an accepted plan.

What the team knows becomes files Claude reads, and the guardrails run as code rather than as habits.

The ask

Marcus picks up the accepted spec. He opens Claude Code in the Spindle repository in , which means Claude can read everything and change nothing. He attaches intent.md and spec.md and asks for an implementation plan: which files change, in what order, and which tests prove it works.

The interview

Claude reads the repository (there is a small starter app), notices there is no shared adapter interface yet, and interviews Marcus. Should the proxy be its own service or a package inside the API? How should a streamed reply be represented internally? Marcus answers, then pushes back the way a senior reviewer would: what could this break, which step is riskiest, and what did you decide not to do? Claude names the streaming layer as the riskiest step, because each provider streams differently, and suggests building the non-streaming path first.

Accepted

They iterate until the plan is something an engineer who never saw the conversation could follow. Marcus saves it as and accepts it. That acceptance is this stage's gate. Because the proxy handles secrets, Linda gives it a quick look too. Plan mode enforces the gate by itself: Claude cannot edit a file until the plan is accepted.

The build

Now Claude builds. Marcus switches to , so Claude applies edits without asking about each one, and he splits the plan into three streams that touch different files. Each stream runs in its own : one session builds the chat window from the Claude Design mock, starting from Claude Design's hand-off prompt, one builds the adapters, and one builds the API and the proxy. Marcus steers all three from one desk, reviewing rather than typing.

What keeps it safe

Three things keep this safe without Marcus watching every keystroke:

  • , a one-page note that every session reads first: how to run and test the app, the conventions, and the mistakes Claude tends to make in this codebase.
  • A skill called provider-adapter that describes exactly how a model adapter must be written, so the OpenRouter adapter and the NVIDIA adapter come out the same shape.
  • : tiny scripts that block reading or editing the local key file, run the formatter after every file change, and refuse anything that looks like an API key. A skill persuades; a hook enforces.

The plan stays true

When the build departs from the plan (it does: the router needs a retry rule nobody planned for), Claude updates plan.md in the same commit, so the plan stays true to the code.

How it flows

Plan first, with Claude in read-only mode. Then build, with hooks guarding every edit.

The old way

An engineer reads the design and starts coding. How the change will be made, down to which files and which tests, stays in their head. The first thing a reviewer sees is the finished code, and by then changing course is expensive. One engineer, one task at a time.

The AI-native way

Work starts with a written plan that Claude produces without touching code. Marcus corrects the plan while it is still a document, then lets Claude build it, often in a single pass, across several parallel sessions. Team knowledge lives in CLAUDE.md and skills; the non-negotiables live in hooks.

Early sign it is working

The share of changes that merge from the first build pass, and the time from plan approval to merged pull request.

Later sign it is working

Rework cycles per change, how often the merged code still matches plan.md, and how often Claude repeats a mistake that CLAUDE.md should have caught.

Stage 04 – Test – Ends by saving the test output and screenshots as evidence

Every session checks its own work before a person sees it.

And the setup that steers Claude gets tested the same way the code does.

The rule

Nothing reaches Marcus until it has already passed a check. That is the rule of this stage, and it is what makes three parallel sessions possible without three reviewers.

The feedback loop

Every Claude session in Spindle has a . The verification block has been in CLAUDE.md since the build kit, and Stage 4 adds the test guard, the done check and the verifier. It says: before you report anything as done, run npm test, npm run build and npm run lint, and paste the output. So the proxy session writes the router, runs the tests, sees one fail (it logged the message text by mistake), fixes the router, runs the tests again, and only then says "done". For the chat window the check is visual: Claude has a browser tool, takes a screenshot of the window, compares it with the approved mock, adjusts the spacing, and repeats. Three rounds is normal.

The verifier

When a session believes the work is finished, a takes over with a fresh memory. It starts the app, sends "hello" through every model in the dropdown, checks that an answer comes back and that the log line contains no message text, and reports what it saw. It is not allowed to fix anything, only to report, so its verdict is not coloured by the assumptions that produced the code.

The first bug

Then a bug turns up: the OpenRouter adapter drops the last word of a streamed GPT reply. Marcus asks Claude to write a failing test that reproduces the bug first, runs it, confirms it fails for the right reason, and commits that test. Only then does Claude fix the adapter. A hook blocks any edit to test files during the fix, so Claude cannot make the test pass by weakening it. A test that existed before the fix, and could not be changed, is the proof the bug is gone.

Testing the setup

Finally the team tests the AI setup itself. Marcus collects twenty real tasks cut from the team's own merged work, such as "add the NVIDIA adapter" or "add a token counter to the model picker", each with the checks that define acceptable: tests pass, lint is clean, no keys in the diff, the provider-adapter skill was followed. This runs after the pipeline on every pull request that changes CLAUDE.md or a skill, and every week; a change to a hook, a workflow or the settings waits for someone to press Run. If a skill change makes the pass rate drop, that change is reviewed before it merges. QA did not disappear; it moved inside the loop.

How it flows

Two loops: the inner one Claude runs on its own work, and the outer one the team runs on Claude's setup.

The old way

The signal that code works arrives late: CI minutes later, a tester days later, production weeks later. When Claude writes the code, a late signal means a person has to check all of it, and that person becomes the bottleneck. QA is a gate at the end.

The AI-native way

Each session verifies its own work and fixes its own mistakes before Marcus looks. The evidence comes from the tools, not from Claude's say-so. And the setup that steers Claude (CLAUDE.md, skills, hooks) has its own regression suite, so a change to a skill cannot quietly make Claude worse.

Early sign it is working

First-pass CI success rate for Claude-written changes, and how long it takes for a production incident to become a permanent eval.

Later sign it is working

Review time per pull request should fall, because the tests catch what reviewers used to catch. Regressions caught in CI versus regressions found in production.

Stage 05 – Deploy – Ends by saving the merged pull request and the release record

Review runs in both directions, and the rules are enforced as Claude acts.

Claude does everything up to the production gate and nothing past it.

The review

Marcus opens a from the server worktree. Within a minute Claude has reviewed it, following , the team's written review policy. Three passes: bugs, security, and whether the change matches spec.md and plan.md. It finds one Important item: the /chat endpoint checks the sign-in but the new /models endpoint does not. It finds two nits about naming. It confirms the code matches the plan, including the retry rule that was added and documented.

The fix

Marcus tags @claude on the Important finding. In the pull request, Claude's own commit adds the fix, and the pull request thread records both the request and the change. The nits he leaves. REVIEW.md caps nits at five so they never drown out the signal.

The approval

Now Linda reviews. She does not read every line. The mechanical evidence is already attached: tests green, lint clean, review findings resolved. She asks the two questions only a person can answer: does this change do what the plan intended, and is the risk acceptable? She approves. The on main accepts changes only through a pull request whose checks pass, and a hook and deny rules stop Claude merging. On a one-account repository GitHub cannot tell Claude's merge from yours, so Linda's approval is written in the pull request before she merges.

The pipeline

The merge runs the 's build and tests. Claude runs inside it too, non-interactively and in a : its CI jobs run on a fresh runner, with a network allow-list and no key files. Spindle's CI token is one person's and lasts a year; a company would use short-lived scoped tokens. When a build fails, a separate triage workflow asks Claude to read the log and say whether the failure looks flaky or real. Deploys run from the team's own machine through three , one per environment.

Going live

Autonomy is tiered by environment. In development, Claude deploys freely. In staging, it deploys and runs the smoke test. In production, it prepares the release and a hook stops it right there: the deploy command needs a release authorization, and only Linda can give one. Linda starts the session with her release authorization for this exact commit, and the gate still asks her to confirm before production. Spindle goes live on the team's own machine: development, staging and production each run locally on their own port. The command was rehearsed in staging that morning, because the next stage may need it with nobody watching (a replayed night).

How it flows

Two human gates: one before the merge, one before production. Everything in between is Claude and the pipeline.

The old way

Review capacity was planned around human output. A pull request waits for a reviewer to read all of it; quality depends on how busy that reviewer is. Deployment and rollback are runbooks a person follows under pressure, and governance happens in review meetings, inconsistently.

The AI-native way

Every pull request gets the same review passes, ranked by severity. Human attention moves up a level, to intent and risk. Rules are enforced by hooks at the moment Claude acts, every time, for everyone. Claude can act all the way to the production gate, and never through it.

Early sign it is working

Time to first review, which should fall to minutes. The share of review comments resolved without a person touching the branch. Time spent waiting at each approval gate.

Later sign it is working

Defects and vulnerabilities caught before merge versus those that escape to production. The standard DORA delivery measures, which scripts/measures.ts reads from the release records and git history.

Stage 06 – Maintain – Ends by saving a new intent.md, with evidence

The loop closes.

A trigger invokes Claude with no person in the path, and what it finds re-enters the pipeline as a new intent.md.

The breach

It is 3:10 a.m. This night is a replay anyone can run with npm run replay: two weeks of normal traffic, then a provider error shape invented for the drill. A small script, not Claude, checks a handful of numbers every five minutes against their normal range: error rate per provider, response time, and the CI test failure rate. Tonight the error rate for GPT replies jumps well outside its .

The diagnosis

The script is deterministic. It does not guess; it looks up what to do in bands.yaml. A small drift only gets logged. Tonight's jump is a 2σ breach, so the script invokes Claude in read-only mode with a narrow set of tools. Claude reads the recent logs and the adapter code and finds the cause: the provider returns an error shape invented for the drill, so our error mapping now treats every reply as "upstream failed". Claude writes its diagnosis as a new intent.md: the problem, the evidence, a proposed outcome, the affected files, and one open question (did the provider change the streaming format too?).

A bigger breach

At 3σ with a deploy in the window, the watcher lets Claude do one more thing: run the pre-approved rollback runbook, with no other tool (Spindle's staging-rollback replay shows it). Any code fix still arrives as a pull request through Stage 5's gate; Claude never has a route around it.

Fix now

At 8:30 a.m. Linda, on call, opens the triage queue and finds the intent waiting with its evidence attached. She picks "fix now". The file enters Stage 1 exactly as Mark's original idea did, and the loop runs again: spec, plan, build, test, review, ship, this time in a few hours rather than a few days. When the fix ships, the team adds an eval for "provider error format changed" to the suite, so the AI setup is tested against this class of bug from now on.

The weekly scan

Once a week, Claude reads the whole repository for vulnerabilities, validates each finding, and proposes patches that go through the same pull request gate as any other change. The hosted, scheduled product, , is part of Claude Enterprise; Spindle's scan is a stand-in, and its setup ships as an example record. A small finding becomes a pull request; a large one becomes an intent.md. The loop no longer has a beginning. It only has gates.

How it flows

Detection is a script and never a model. How far Claude may go depends on how far the number has drifted.

The old way

Maintenance is reactive. An alert fires at 3 a.m. and may be missed. A ticket sits in the backlog until someone picks it up. Post-mortem actions may never reach the codebase if another fire starts first. A person has to notice, then start the whole process from scratch.

The AI-native way

A breached band, a ticket, or a schedule invokes Claude without a person in the path. Claude diagnoses, acts only through gated routes, and writes an intent.md that flows through the stages above. People triage and review the work; they no longer have to start it.

Early sign it is working

Time from a breached band to an intent.md in the triage queue, against the old time from incident to post-mortem action. Share of repositories on a scheduled security scan.

Later sign it is working

The share of findings that become merged fixes, and repeat incidents of the same kind, which should fall as each fix adds an eval to the suite.

After the loop

What actually changed.

Six stages, six files, six gates. The stage names are the same as in the traditional lifecycle. What changed is who starts the work, where the team's knowledge lives, and where people spend their attention: at the gates, reviewing what Claude flagged, instead of starting every stage from a blank page.

StageEnds by savingThe gate (a person decides)In the Spindle story
1 Planintent.mdRahul accepts or closes the intentMark's "one chat window, any model" idea, in his own words
2 Designspec.md and the screen mockRahul accepts the spec; Linda joins for higher riskSix requirements, four packages, three concerns resolved before engineering saw it
3 Buildplan.md, then code and testsMarcus accepts the plan before any code is writtenThree worktrees at once; hooks guard the key file
4 TestTest output and screenshots, as evidenceA person reviews if the eval pass rate dropsA verifier sends "hello" through every model; 20 evals run weekly and on setup changes
5 DeployThe merged pull request and the release recordLinda approves the merge, then authorizes productionClaude found the missing sign-in guard; a hook held the production door
6 MaintainA new intent.md with evidenceThe on-call engineer triages: fix, schedule, or dismissA replayed 3:10 breach became a triaged intent

What Claude never does

Approve its own work. The approval is a person's written decision in the pull request.
Merge to the main branch. A hook and deny rules stop it, and a person merges.
Pass the production gate. The hook needs a named authorization.
Edit a test while fixing a bug. A hook blocks it.
See an API key. Keys come from a local file outside the project at runtime; a deny rule and a hook stop Claude reading it.
Decide whether a metric breach matters. A plain script decides; Claude is invoked afterwards.

Every dotted term, in one place

Plain-English glossary.

The same definitions the dotted terms show when you rest on them, listed alphabetically so you can come back to one without hunting for it in the story.