Workshop Studio
participantPublic visitor

TDD with Claude

You have a validated spec, an architecture plan, and CLAUDE.md pointing to the active work. Implementation is plan execution — not creative exploration. The approach: convert acceptance criteria into tests BEFORE writing code, then implement to pass them.

Acceptance Criteria → Phase 1 → Generate Tests (Tests fail — nothing exists yet) → Implement (Build to pass the tests) → Tests Pass (Every criterion verified)

This gives Claude two constraints: the spec/architecture (design intent) AND the failing tests (executable expectations). Double verification — and a natural guardrail against over-engineering.

SignalHealthyUnhealthy
Test generationTests map 1:1 to acceptance criteriaTests are generic or don't trace to requirements
Exploration phaseClaude reads 3-5 files before planningClaude starts writing code immediately without reading
Pattern matchingPlan follows existing conventionsPlan introduces different patterns than the rest of the project
Scope adherencePlan covers exactly what the spec saysPlan adds extra features 'while I'm at it'
Error formatPlan references the architecture's error formatPlan invents a new error shape
Test resultsTests fail before implementation, pass afterTests pass before implementation (tests are wrong)

You have all Phase 1 and Phase 2 artifacts in place. Start with a clean context — your artifacts are the context bridge, not chat history.

Step 1: Start fresh context

Clear your conversation history to start fresh for implementation. Your Phase 2 artifacts (spec, architecture plan, CLAUDE.md) are already in files — you don't need the chat history from previous phases:

Run in terminal:

1
/clear

Observe: Context is cleared. Claude will re-read CLAUDE.md automatically, which now points to the active work (spec, architecture plan). All your Phase 1 and 2 artifacts are in files — nothing was lost. This is why we write decisions to files, not chat.


Step 2: Generate tests from acceptance criteria

First, generate tests from your acceptance criteria. No implementation yet — just the tests:

Tell Claude:

Read the acceptance criteria and API contract in
docs/specs/task-comments.md.

Write tests that verify EVERY acceptance criterion:
- Positive tests (valid input accepted)
- Negative tests (invalid input rejected, missing fields, non-existent tasks)
- Assert the API contract exactly (status codes, response shape, error format)
- Verify behavior matches docs/architecture/task-comments.md

Follow the project's existing test structure. If none exists,
use the ecosystem default.
Do not implement anything yet — only write the tests.

Observe: Claude reads the validated spec and generates test cases that map to each criterion. Check: does every acceptance criterion have at least one test? Are the assertions specific (checking status codes, error format, field names) rather than vague? The spec is the build contract — not the raw stories from Phase 1 — because validate-spec tightened its criteria and added the API contract the tests need to assert against.


Step 3: Confirm red — all tests fail

Run the tests. Every single one should fail — that is the point:

Tell Claude:

Run the tests.

Observe: Every test fails. This is the 'red' in red-green-refactor. Check three things: (1) the test runner found all test files — no configuration issues, (2) every test fails because the feature is not implemented, not because of syntax errors or import problems, (3) count the tests — this is your progress meter for the implementation phase.


Step 4: Claude writes implementation plan

Now ask Claude to write an implementation plan scoped to the full spec, not just the tests:

Tell Claude:

Now implement the ACTIVE unit. Every acceptance criterion in the
spec must be implemented — including [UI] criteria that aren't
covered by automated tests. The tests are a subset of the contract,
not the whole contract.

Read these artifacts first:
- Feature spec: docs/specs/task-comments.md
- Architecture: docs/architecture/task-comments.md

Before writing any code:
1. Write an implementation plan to docs/plans/task-comments.md
   with checkboxes for each step
2. Flag any decisions where you need my input (mark with QUESTION:)

Do not implement yet — just write the plan for my review.

Observe: Claude reads the spec, architecture, and existing code, then writes an implementation plan. Review the plan: does the file order make sense? Does it cover every acceptance criterion — including the [UI] ones that don't have tests? Are there scope creep items to remove?


Step 5: Resolve open questions

If the plan has QUESTION: markers, ask Claude to resolve them with its best recommendation rather than leaving them as open questions for you:

Tell Claude:

For each QUESTION in the plan, replace it with your
best recommendation based on the spec and architecture.
Explain your reasoning briefly inline.
Update the plan file.

Observe: Claude replaces each QUESTION with a concrete recommendation and a one-line rationale. Now you review a complete plan with proposed answers, not a plan with blanks to fill. Much easier to approve or override a recommendation than to answer from scratch.


Step 6: Handoff and execute

Open docs/plans/task-comments.md in your editor. Review it, then hand it back:

Tell Claude:

I reviewed the plan in docs/plans/task-comments.md.
Follow it as specified. Implement one step at a time,
mark each checkbox as done. When all steps are complete,
run the tests.

Observe: Claude implements step by step, running tests periodically. Watch the test count: 0/N passing at start, then climbing as each piece lands. When all tests pass, implementation is complete. Check that error responses match the architecture's error format.


Step 7: Verify traceability

Verify all tests pass and check traceability:

Tell Claude:

Run all tests. For each acceptance criterion on the task
comments stories in requirements/stories.md, point to the
test that asserts it so I can see the chain
stories → spec → tests is intact.

Do not modify tests or code.

Observe: This is the end-to-end traceability check. stories.md is the original user-facing contract; the spec refined it; the tests execute it. Seeing each criterion mapped to a specific test confirms the chain — not just "tests match the spec" but "tests match what the user actually asked for." This step only surfaces the mapping; it does not change anything.


Check the plan: did Claude tick boxes it could not actually verify?

Open docs/plans/task-comments.md and scan the checklist. Depending on the model you are running, it has been observed that Claude sometimes designates UI acceptance criteria as manual verification steps — because the project has no browser-driving harness wired up to test them automatically — and then ticks those manual items anyway as part of finishing the implementation loop. Stronger models are less prone to this; weaker ones fall into it more often. A [ ] in the plan has no format distinction between "I ran this" and "a human will run this," so once the loop enters iteration mode every unchecked box looks like progress. The app works in your browser today because Claude wrote reasonable code from a careful spec, not necessarily because anything verified the UI contract end-to-end.

This is the exact failure mode Phase 4 addresses structurally. The Ralph Loop replaces the manual checklist with a three-layer binary gate — ESLint, Jest, and Playwright driving a real browser. UI acceptance criteria become tests that either pass or fail — no self-reported [x]. When you see Ralph run in the next chapter, this is the problem it is solving: an implementation loop cannot lie about behaviour it was structurally unable to check.

Step 8: Manual verification

Open TaskFlow in your browser, click a task card to open the detail panel, and try posting a comment. Confirm it appears in the list without the page reloading, that the textarea clears after submit, and that an empty submit does nothing. This is your sanity check before the self-review — the part Claude could not verify for itself.

Key Insight

You just closed the artifact chain: human intent (Phase 1 requirements) → acceptance criteria (Phase 1 stories) → executable tests → implementation code. Every link is traceable. The tests existed before the code — they defined what 'done' means. Claude implemented to satisfy them, not the other way around.