TDD with Claude
You have a validated spec, an architecture plan, and CLAUDE.md pointing to the active work. Implementation is plan execution — not creative exploration. The approach: convert acceptance criteria into tests BEFORE writing code, then implement to pass them.
Acceptance Criteria → Phase 1 → Generate Tests (Tests fail — nothing exists yet) → Implement (Build to pass the tests) → Tests Pass (Every criterion verified)
This gives Claude two constraints: the spec/architecture (design intent) AND the failing tests (executable expectations). Double verification — and a natural guardrail against over-engineering.
| Signal | Healthy | Unhealthy |
|---|---|---|
| Test generation | Tests map 1:1 to acceptance criteria | Tests are generic or don't trace to requirements |
| Exploration phase | Claude reads 3-5 files before planning | Claude starts writing code immediately without reading |
| Pattern matching | Plan follows existing conventions | Plan introduces different patterns than the rest of the project |
| Scope adherence | Plan covers exactly what the spec says | Plan adds extra features 'while I'm at it' |
| Error format | Plan references the architecture's error format | Plan invents a new error shape |
| Test results | Tests fail before implementation, pass after | Tests pass before implementation (tests are wrong) |
You have all Phase 1 and Phase 2 artifacts in place. Start with a clean context — your artifacts are the context bridge, not chat history.
Step 1: Start fresh context
Clear your conversation history to start fresh for implementation. Your Phase 2 artifacts (spec, architecture plan, CLAUDE.md) are already in files — you don't need the chat history from previous phases:
Run in terminal:
1
/clearObserve: Context is cleared. Claude will re-read CLAUDE.md automatically, which now points to the active work (spec, architecture plan). All your Phase 1 and 2 artifacts are in files — nothing was lost. This is why we write decisions to files, not chat.
Step 2: Generate tests from acceptance criteria
First, generate tests from your acceptance criteria. No implementation yet — just the tests:
Tell Claude:
Read the acceptance criteria and API contract in
docs/specs/task-comments.md.
Write tests that verify EVERY acceptance criterion:
- Positive tests (valid input accepted)
- Negative tests (invalid input rejected, missing fields, non-existent tasks)
- Assert the API contract exactly (status codes, response shape, error format)
- Verify behavior matches docs/architecture/task-comments.md
Follow the project's existing test structure. If none exists,
use the ecosystem default.
Do not implement anything yet — only write the tests.Observe: Claude reads the validated spec and generates test cases that map to each criterion. Check: does every acceptance criterion have at least one test? Are the assertions specific (checking status codes, error format, field names) rather than vague? The spec is the build contract — not the raw stories from Phase 1 — because validate-spec tightened its criteria and added the API contract the tests need to assert against.
Step 3: Confirm red — all tests fail
Run the tests. Every single one should fail — that is the point:
Tell Claude:
Run the tests.Observe: Every test fails. This is the 'red' in red-green-refactor. Check three things: (1) the test runner found all test files — no configuration issues, (2) every test fails because the feature is not implemented, not because of syntax errors or import problems, (3) count the tests — this is your progress meter for the implementation phase.
Step 4: Claude writes implementation plan
Now ask Claude to write an implementation plan scoped to the full spec, not just the tests:
Tell Claude:
Now implement the ACTIVE unit. Every acceptance criterion in the
spec must be implemented — including [UI] criteria that aren't
covered by automated tests. The tests are a subset of the contract,
not the whole contract.
Read these artifacts first:
- Feature spec: docs/specs/task-comments.md
- Architecture: docs/architecture/task-comments.md
Before writing any code:
1. Write an implementation plan to docs/plans/task-comments.md
with checkboxes for each step
2. Flag any decisions where you need my input (mark with QUESTION:)
Do not implement yet — just write the plan for my review.Observe: Claude reads the spec, architecture, and existing code, then writes an implementation plan. Review the plan: does the file order make sense? Does it cover every acceptance criterion — including the [UI] ones that don't have tests? Are there scope creep items to remove?
Step 5: Resolve open questions
If the plan has QUESTION: markers, ask Claude to resolve them with its best recommendation rather than leaving them as open questions for you:
Tell Claude:
For each QUESTION in the plan, replace it with your
best recommendation based on the spec and architecture.
Explain your reasoning briefly inline.
Update the plan file.Observe: Claude replaces each QUESTION with a concrete recommendation and a one-line rationale. Now you review a complete plan with proposed answers, not a plan with blanks to fill. Much easier to approve or override a recommendation than to answer from scratch.
Step 6: Handoff and execute
Open docs/plans/task-comments.md in your editor. Review it, then hand it back:
Tell Claude:
I reviewed the plan in docs/plans/task-comments.md.
Follow it as specified. Implement one step at a time,
mark each checkbox as done. When all steps are complete,
run the tests.Observe: Claude implements step by step, running tests periodically. Watch the test count: 0/N passing at start, then climbing as each piece lands. When all tests pass, implementation is complete. Check that error responses match the architecture's error format.
Step 7: Verify traceability
Verify all tests pass and check traceability:
Tell Claude:
Run all tests. For each acceptance criterion on the task
comments stories in requirements/stories.md, point to the
test that asserts it so I can see the chain
stories → spec → tests is intact.
Do not modify tests or code.Observe: This is the end-to-end traceability check. stories.md is the original user-facing contract; the spec refined it; the tests execute it. Seeing each criterion mapped to a specific test confirms the chain — not just "tests match the spec" but "tests match what the user actually asked for." This step only surfaces the mapping; it does not change anything.
Step 8: Manual verification
Open TaskFlow in your browser, click a task card to open the detail panel, and try posting a comment. Confirm it appears in the list without the page reloading, that the textarea clears after submit, and that an empty submit does nothing. This is your sanity check before the self-review — the part Claude could not verify for itself.