A spec-driven code generation system for greenfield projects.
The spec is the source of truth. Code is a build artifact. Convergence and determinism do not come from the prose — natural language under-constrains, and two readers (or two model runs) will fill the gaps differently. They come from a machine-checkable contract layer that sits between the prose and the code.
The prose says what we want. The contracts say it in a form that collapses the space of acceptable outputs. The tests, generated from the contracts, are the real interface: they pin observable behavior. Code is accepted if and only if it passes those tests — its internal structure is explicitly nobody's business.
This inverts the usual relationship. Normally code is primary and docs rot beside it. Here the spec is primary and code is regenerable. You edit the spec; the pipeline tells you which behavioral promises changed and re-derives the tests that enforce them.
Intent /spec/*.md human-editable markdown; the WHAT and WHY
│ /spec/README.md carries YAML frontmatter = the machine handle
▼
Contracts /contracts extracted from frontmatter:
│ · data schemas (JSON Schema)
│ · interface signatures (typed stubs)
│ · acceptance criteria
▼
Tests /tests generated from contracts:
│ · property-based + acceptance tests
│ · the real, enforced interface to behavior
▼
Code /src generated to satisfy the tests.
structure unconstrained; only behavior is.
-
Intent — Human-editable markdown describing the system, in
/spec/*.md. The root/spec/README.mdcarries structured YAML frontmatter that is the actual machine-readable handle. Prose is for motivation; frontmatter is for the build. -
Contracts — Extracted from the frontmatter: type signatures / API schemas, data schemas (JSON Schema today; SQL DDL or protobuf are natural future targets), and acceptance criteria. These narrow the output space far more than English can.
-
Tests — Generated from the contracts. Property-based and acceptance tests are the real interface that pins observable behavior. Code is accepted only if it passes them.
-
Code — Generated to satisfy the tests. Structural form is not constrained; only behavior is. Any implementation that passes is correct.
The point of building this is incremental regeneration. Editing one requirement in the markdown should regenerate only the affected acceptance tests, producing a small, reviewable diff — not a wholesale rebuild you can't read.
This works because every acceptance criterion has a stable id and a
content hash. The diff pipeline compares two versions of the spec, reports
which criteria were added / removed / changed, and regenerates exactly the
matching tests. See spec/SCHEMA.md for the full diff contract.
| Path | Role |
|---|---|
/spec |
Intent. SCHEMA.md (frontmatter format) + README.md (the reference spec). |
/contracts |
Extracted contracts: data/*.schema.json, interfaces/*. |
/tests |
Generated tests — the enforced behavioral interface. |
/src |
Generated implementation. Unconstrained in form. |
/pipeline |
The tooling: extraction, test generation, and the spec-diff engine. |
tinylink, a small URL shortener, is the worked example — small enough to fully
specify, rich enough to have an entity, a three-operation API, and real edge
cases (invalid input, unknown codes, observe-without-mutate). It lives entirely
in the frontmatter of /spec/README.md.
Python for the pipeline itself (fast iteration, mature property-testing via Hypothesis, first-class JSON Schema tooling). The generated code can target any language — only the test interface is contractual.
Early. Tasks completed so far:
- Directory structure (
/spec,/contracts,/tests,/src,/pipeline) - Frontmatter schema (
spec/SCHEMA.md) - Reference spec (
spec/README.md, thetinylinkproject) - Extraction: frontmatter → contracts
- Test generation: contracts + acceptance → executable tests
- Spec-diff → test-diff pipeline
Extraction, test generation, and the diff pipeline are deferred pending review of the schema and reference spec.