Spec Kit on a project that already exists: the real flow, step by step
By Seyed Masoud Hosseini · · Engineering
My first question after installing GitHub's Spec Kit was "do I start with /speckit-plan?" The answer is no. What Spec Kit is, the order of commands for adding a feature to a codebase you already have, what it cost on a real feature on this site, and the three bugs it didn't catch.
The first thing I asked after installing GitHub's Spec Kit was: "I want to add a new feature. Should I run /speckit-plan?"
No. It's the most natural guess, and it's wrong. This post is the answer I wish I'd had at that moment: what Spec Kit is, what each command is for, and the order to run them when you're adding to a project that already exists.
It isn't a summary of the docs. I set Spec Kit up on this website and used it to build one real feature, start to finish, in a day: a table of contents that follows you as you read an article. So alongside the steps you'll find what each one actually produced, what the whole thing cost, and what it missed.
If you just want the steps, skip to the flow. If you want to know whether it's worth it, skip to what it cost.
What Spec Kit is, in one minute
When you ask an AI coding agent (Claude Code, Copilot, Gemini and others) to "add favourites to the app", it guesses. It guesses what you meant, which files to touch, which libraries to use and when it's done. On a small script that's fine. On a real codebase you get code that works but ignores your conventions, adds a dependency you didn't want, or solves a slightly different problem.
Spec Kit is GitHub's open-source answer to that. The idea is called spec-driven development: before any code is written, you and the agent agree in writing on what you're building and why, then how, then the list of steps. Only then does the agent write code, and it works from those documents rather than from a one-line prompt.
In practice it's two things:
- A small command-line tool,
specify, that you run once to add Spec Kit to a project. - A set of slash commands inside your agent (
/speckit-specify,/speckit-planand so on). Each one writes or checks a Markdown file in your repo.
Because everything ends up as plain files in the repo, you can read them, fix them, commit them and review them like code.
One naming note: in Claude Code the commands are skills and use a dash (/speckit-plan). Many other agents, and most of the official docs, use a dot (/speckit.plan). Same commands.
The five documents
Everything Spec Kit does comes down to five kinds of file. Once you know them, the commands make sense.
| Document | Answers | Written by | How often |
|---|---|---|---|
constitution.md | What rules does this project always follow? | /speckit-constitution | Once per project |
spec.md | What are we building, for whom, and why? | /speckit-specify | Once per feature |
plan.md (plus research.md, data-model.md and friends) | How will we build it in this codebase? | /speckit-plan | Once per feature |
tasks.md | What are the steps, in order? | /speckit-tasks | Once per feature |
| The code | — | /speckit-implement | Once per feature |
The constitution lives in .specify/memory/. Everything for a feature lives in its own numbered folder, like specs/001-article-toc/.
The key idea is that each document is built from the one before it. The plan is written from the spec, the tasks from the plan, the code from the tasks. That's why you can't start with /speckit-plan: there's nothing yet for it to plan from.
The order, and the one rule
For every new feature:
/speckit-specify "what you want"
/speckit-clarify (optional, recommended)
/speckit-plan
/speckit-tasks
/speckit-analyze (optional, recommended)
/speckit-implement
/speckit-converge (repeat implement → converge until it says done)
The one rule: what and why before how. Specify is about the reader or user and the behaviour, with no technology in it. Plan is where frameworks, files and database tables come in.
The optional ones:
/speckit-clarifyasks you up to five pointed questions about gaps in the spec and writes your answers back into it. It's the cheapest place in the whole process to catch a misunderstanding./speckit-analyzereads spec, plan and tasks together and reports anything that contradicts or is missing. It doesn't change anything./speckit-checklistgenerates an extra review checklist for a specific concern (accessibility, security, and so on). It's an extra gate you can add, not a required step before planning, whatever some summaries say.
/speckit-converge is newer and easy to miss. After implementing, it compares the code with the spec, plan and tasks, and appends any unfinished work to tasks.md. You run implement again, then converge again, until it reports the feature complete.
And one command you don't repeat: /speckit-constitution is set up once per project, not once per feature.
Why an existing project is different
On a brand-new project, the agent can choose the stack. On an existing one it mustn't. You already have a framework, a folder structure, a test setup, a way of naming things, and probably some opinions you've never written down.
A plain install doesn't know any of that. It gives you generic templates. So on an existing project, the constitution does the most important work: it's where you pin down the stack and the conventions, and tell the agent to reuse what's there before creating anything new. Get that right and every later step inherits it.
Two things the official guide for existing projects says that I'd underline:
- Don't specify what already exists. Your old code stays as it is and acts as context. The first spec is for the next change, not a retroactive description of the whole system.
- Pick a small, bounded first feature. Something you can review on its own in a day, not "document the architecture".
The flow on a real project
This site is a good test case: a personal site that's been growing for months, with its own build scripts, four languages, React without any framework on top, and a few strict rules (for example, nothing loads from outside CDNs). The feature: a table of contents for articles, beside the text on desktop and as a small control on phones.
Step 1. Install Spec Kit (once)
You need uv (a Python tool installer) and git. Then:
uv tool install specify-cli
Start from a clean working tree, so that afterwards git diff shows exactly what Spec Kit added:
cd your-project
git checkout main && git pull
git status # commit or stash anything open first
specify init --here --force --integration claude
--here means "this folder". --force lets it write into a folder that already has code in it. It doesn't touch your application files, but it can overwrite files at paths Spec Kit manages, which is why the clean tree matters. Swap claude for your agent (copilot, gemini, codex and more are supported).
It adds a .specify/ folder (templates, scripts, the constitution) and the commands for your agent. In Claude Code they show up as skills in .claude/skills/. Look over the diff, then commit it on its own.
A detail that caught me out: /speckit-specify does not create a git branch. Branches are handled by an optional git extension. My first spec landed in specs/001-article-toc/ while I was still on main. If you want one branch per feature (and on a team you do), create it yourself before specifying, or add the extension:
specify extension add git
Step 2. Write the constitution from the existing code (once)
This is the brownfield step. Don't write your principles from scratch, and don't invent standards to fill the template. Tell the agent to read the codebase (README, contribution notes, CI config, tests) and write down what's already true:
/speckit-constitution Read the existing codebase and derive principles
from its current stack, structure, conventions and testing approach.
On a codebase with a database or an API, name the non-negotiables too: "the stack is fixed, no new libraries without approval, reuse existing services before creating new ones, don't refactor unrelated code".
For this site, it came back with five principles, all things I had been doing but never written down in one place:
- Content lives as files, and the site is build output. No database, no CMS.
- No framework and no dependencies in the tooling. Build scripts use Node built-ins only, and any new npm package needs a written reason.
- Self-hosted and private by default. No scripts, fonts or images from outside CDNs.
- Every page is crawlable and works in all four languages.
- Logic is tested, and
lint,typecheckandtestmust pass before any commit.
It also added a rule that every plan must be checked against these five, and that any exception has to be written down with the simpler option that was rejected.
Now read it properly. This is the one document where a mistake follows you into every feature. Mine contained one rule that the code was already breaking: some page metadata still said something the constitution said pages shouldn't. The agent pointed it out; it would have been easy to miss. Fix what's wrong by hand, then commit.
Step 3. Specify the feature: what and why
Describe the feature the way you'd explain it to a person. No technology. Mine was one rough sentence, typos included:
/speckit-specify on articles i want to show list of heading of that article
to user on desktop show on side and on mobile show on top or buttom
From that it wrote a 200-line spec.md with:
- Context: what already exists. It noticed articles already have an "On this page" box that scrolls away as soon as you start reading, and framed the feature as fixing that. I hadn't mentioned it.
- User stories, each with a priority and a way to test it on its own: the desktop side list (P1), the phone control (P2), and working for every reader and language (P3).
- Acceptance scenarios in Given / When / Then form, such as "given the reader has scrolled halfway, the list is still on screen and the current section is highlighted".
- Edge cases, requirements, success criteria and assumptions.
- A checklist in
checklists/requirements.mdconfirming there's no implementation detail in the spec.
It also stopped and asked me the one thing my sentence left open: top or bottom of the screen on phones? It laid out the options with their trade-offs and recommended the bottom (within thumb reach, and clear of the menu at the top). I agreed.
Your job here is to read it and correct it. A wrong sentence in the spec costs you ten seconds now and a wrong feature later.
Step 4. Clarify
/speckit-clarify
It looks for decisions the spec still leaves open and asks about them one at a time, as multiple choice with a recommended answer. Mine asked three:
- On desktop, should the list sit to the left or the right of the text? (I chose left, against its recommendation of right.)
- On phones, should the bottom control stay on screen, or hide while you scroll down and come back when you scroll up, like the site menu? (Hide, like the menu.)
- On phones, should the existing "On this page" box stay at the top, with the bottom control appearing only once it has scrolled away? (Yes.)
The third one mattered most. The first draft of the spec quietly removed that box everywhere, which would have left phone readers with no outline at the start of an article. Each answer went into a Clarifications section of the spec and into the requirements it changed, so the decisions are recorded, not just remembered.
Step 5. Plan: how, in your real code
Now the technology comes in. You can run it bare, or point it at code to reuse:
/speckit-plan Build on the existing "On this page" list rather than
replacing it. No new dependencies. Put the logic in pure functions with tests.
It writes plan.md and, depending on the feature, research.md (the design decisions and the options it rejected), data-model.md, contracts/ and quickstart.md (how to check it works). The plan includes a Constitution Check against your principles.
What to check: does the plan name your real files and patterns, or invented ones? Mine did name the real ones: three existing files to change, no new files, no new dependencies, the logic that doesn't need a browser as small pure functions in an existing module with tests beside the existing ones, and the existing "On this page" translation reused so there was nothing new to translate. If yours mentions a component, a folder or a library that doesn't exist in your repo, it hasn't looked properly. Say so and run it again.
Step 6. Tasks, then analyze
/speckit-tasks
/speckit-analyze
/speckit-tasks turns the plan into a numbered checklist in tasks.md, grouped into phases by user story, with tasks that can run in parallel marked. Mine had 23 tasks. Because the stories were prioritised, the desktop list came first and formed a working first version on its own.
One thing I liked: the spec never asked for tests, but the constitution requires them for new logic, so the tasks included writing the tests first. That's the constitution doing its job.
/speckit-analyze then reads spec, plan and tasks side by side and reports problems: a requirement with no task, a task that contradicts the spec, something that breaks the constitution. I skipped it on this feature, because the three documents had been written minutes apart and I'd read each one. On anything larger, or when several people touched the documents, run it and fix what it flags before you write any code.
Step 7. Implement, then converge
/speckit-implement
/speckit-converge
The agent works through tasks.md in order and ticks off each task as it's done. For anything bigger than a small feature, don't let it run the whole list in one go:
/speckit-implement phases 1-2 only
Review, run the app, then continue. Then run /speckit-converge: if anything in the spec isn't built yet, it adds the missing work to tasks.md, and you implement again.
Step 8. Run it yourself, then ship
This is the step no document replaces. Open the thing in a browser, on a phone, in the other languages, in dark mode. Then the usual:
- run the tests (here:
npm run build,npm run lint,npm run typecheck,npm test) - read the diff yourself
- commit the
specs/001-…folder together with the code, so the reasoning lives next to the change - open a PR, review, merge
What it didn't catch
All 23 tasks were done and every test passed. Then I opened the page in a browser and found three bugs. None of them was in the spec, the plan or the tests, because none of them could be seen without a real page:
- The wrong section was highlighted while scrolling. Each paragraph on this site slides in slightly as it appears. Measuring where a heading was during that animation put it 30–70 pixels lower than its real position. The fix was to measure where the heading actually sits, not where it's drawn.
- "Reduce motion" still scrolled smoothly. The code asked the browser for a normal jump for readers who've turned on reduced motion, but the site's CSS makes every scroll smooth by default, so the "normal" jump still animated. It had to ask for an instant jump explicitly.
- The highlight was one section behind on very small phones. At 320 pixels wide, the site menu is slightly too wide, so the phone zooms the page out, and the visible screen no longer starts at the top of the page. That menu problem was older than the feature.
The lesson isn't that Spec Kit failed. The spec was right about what should happen. It's that a spec is a promise about behaviour, and only running the thing checks the promise. The quickstart.md the plan wrote is what led me to these bugs: it listed exactly what to try in the browser. So write that file, then actually do what it says.
What it cost
Here's the honest accounting for one small feature:
- 779 lines of Markdown (spec, plan, research, data model, contract, quickstart, tasks, checklist), plus a 134-line constitution written once.
- 255 lines of code changed, across seven files, including 31 lines of tests.
- Four questions answered by me, about the things I actually cared about.
That's three lines of documents for every line of code. Birgitta Böckeler, who compared Spec Kit with two similar tools on Martin Fowler's site, put the downside plainly: it created "a LOT of markdown files" that were "repetitive, both with each other, and with the code", and she'd "rather review code than all these markdown files". She also saw the agent ignore what the documents said. Fair on both counts: nothing checks that the code matches the spec, except you, the tests and converge.
What I got for the cost: four real decisions made before any code, instead of discovered in review; a plan that stayed inside the site's rules without me repeating them; and a written record of why the feature works the way it does, next to the code.
When to use it, and when not
Use it for a feature that touches several parts of the code, has real choices in it (layout, behaviour, edge cases), or will be read by someone else later. That's where the questions pay for the documents.
Don't use it for a typo, a one-line fix, or a change you could describe completely in one sentence. Böckeler watched a similar tool turn a small bug fix into four user stories with sixteen acceptance criteria. That's a sledgehammer. For bugs, Spec Kit has a separate, lighter bug-fix workflow, available as an extension, that doesn't go through a full spec.
Decide how your specs will age
One question Spec Kit leaves to you: what happens to specs/001-article-toc/ when the feature changes next year? The docs describe three options:
- Flow-forward: finished feature folders are history. A change becomes a new feature,
002-…, that refers back to the old one. - Living spec:
spec.mdis the contract. You edit it first and regenerate the plan and tasks from it. - Flow-back: edit whichever document the discovery lands in, then bring the others back in line.
None is the default. Pick one and write it in the constitution, so the next person knows which document to trust. For a personal site, flow-forward is the least work.
The next feature
From an updated main, start again at step 3. The next feature becomes 002-…. Don't install again and don't redo the constitution. Change the constitution only when a rule actually changes, by running /speckit-constitution again with the change.
Mistakes to avoid
- Starting with
/speckit-plan. No spec, nothing to plan from. Always specify first. - Putting technology in the spec. "Use a React hook" belongs in the plan. The spec says what the reader sees and does.
- A generic constitution on an existing project. If it doesn't name your stack and your rules, the agent will bring in its own.
- Not reading the generated files. The point is that you check each document before the next step. Skip that and you're back to guessing, just slower.
- Treating "all tasks done" as done. Run converge, then run the app. Mine had three bugs after the last box was ticked.
- Using it for everything. Small fixes don't need four documents.
Cheat sheet
Once per project
uv tool install specify-cli
specify init --here --force --integration claude (clean tree first)
specify extension add git (if you want a branch per feature)
/speckit-constitution derive principles from the existing code
→ review .specify/memory/constitution.md, commit
Every feature
/speckit-specify what + why, no tech → specs/NNN-name/spec.md
/speckit-clarify answer the questions → written back into spec.md
/speckit-plan how, in your real code → plan.md, research.md…
/speckit-tasks → tasks.md
/speckit-analyze fix what it flags (read-only report)
/speckit-implement phases 1-2 only, review, continue
/speckit-converge repeat with implement until it reports complete
→ run the app yourself, tests, diff, PR, merge. Next feature: back to specify.
It feels like a lot of documents the first time, and it is. By the second feature it's just the order you work in: what, then how, then steps, then code, reading each one before moving on. Then you open the browser anyway.