# Why AI-Generated Tests Can’t Catch Their Own Bugs — And What Comes Next
## The Measurement That Changed the Conversation
A recent study from a major technology company demonstrated that when developers write formal specifications before generating tests, bug detection rates improve by nearly 10 percentage points. This finding landed with significant weight in the software engineering community, but it also raised an uncomfortable question that has occupied researchers and practitioners alike: *where does the specification come from?*
In the study under discussion, the specification was extracted by reading the existing code. The agent analyzed the implementation, documented its assumptions about pre-conditions, post-conditions, and edge cases, and then used that documentation as the basis for generating test cases. The results were impressive — but they also revealed a fundamental limitation that many had suspected all along.
## The Three Eras of Software Testing
To understand why this matters, it helps to trace how software testing has evolved over the past two decades. Each era shifted the *source of truth* for what software should do further upstream — away from the code itself and toward earlier, more human-readable artifacts.
### TDD: Tests as the Specification
Test-Driven Development established that writing tests before writing code produces better designs. The principle is straightforward: describe the expected behavior first, then make it pass. The approach works, but it leaves the specification embedded in a form that only programmers can read fluently — executable code with assertions. Product managers, business stakeholders, and domain experts are effectively locked out of the conversation.
### BDD: Bridging the Language Gap
Behaviour-Driven Development addressed this problem by introducing structured natural-language scenarios, most commonly written in Gherkin syntax using “Given, When, Then” patterns. Before a single line of code is written, business analysts, developers, and testers collaborate around the same artifact. They negotiate what “correct behavior” looks like in language everyone can understand and agree upon.
The shared specification became the contract between what the software should do and what it actually does.
### SDD: The Machine-Readable Contract
Specification-Driven Development represents the latest evolution, arriving hand-in-hand with AI coding agents. In this model, a formal specification — often written in structured notation like EARS (Easy Approach to Requirements Syntax) — serves as the machine-readable contract. AI agents can consume this spec to plan tasks, generate tests, and produce implementation code.
The specification stops being a side note to the development process and becomes the actual engine that drives it.
This progression moved the source of truth further and further from the code itself — a clear and positive direction. But there is a critical piece missing, and it is the piece this article explores.
## The Independence Problem
Here is the core issue that has driven much of the recent debate: when the same AI agent both writes the specification and writes the code, something deeply problematic can happen.
### Divide and Conquer — But Along the Wrong Line
Modern SDD workflows divide labor efficiently. One agent reads the specification, breaks the work into tasks, generates test cases, and another agent writes the implementation. On the surface, this looks like a textbook application of divide-and-conquer strategy.
But divide-and-conquer only works when the cut you make separates the failure path. If the specification and the code both come from the same reading of the same ambiguous sentence, the failure runs through the *knowledge base*, not through the *work pipeline*. Every agent shares the same assumptions, the same blind spots, and the same unresolved ambiguities.
### The Contract Must Come From Outside the Code
A contract, properly understood, is a statement of what the code *must* do — written before the code exists and never derived from it. When an AI agent generates its contract by reading the implementation it is about to write, it is documenting its own decisions. The contract becomes a mirror, not a measuring stick.
This is the difference between two things that share a name but have opposite provenance:
– **A contract read from code**: tells you what the code decided to do.
– **A contract written before code**: tells you what the code was supposed to do.
The first can be useful for documentation and reasoning. The second is the only thing that can tell you when the code is wrong.
## What the Experiment Reveals
The author of the open-source library at the center of this discussion has been running controlled experiments to isolate the variable that Google’s study could not. The setup is simple in principle, though painstaking in execution:
1. **One agent** receives only the requirements — the decisions that have been made about what the system should do.
2. **That agent writes the implementation**, making whatever judgment calls it needs to resolve ambiguities.
3. **A second agent** receives only the acceptance criteria — the consequences that follow from those decisions — which the first agent never sees.
4. **The second agent generates the test suite** from the acceptance criteria alone.
5. **The test suite runs against the implementation.**
Across twenty runs on two distinct tasks, the test suite disagreed with the code every single time — ten out of ten on each task. By “disagreed,” the author means a concrete test failure: the suite asserted one behavior, and the implementation produced another.
What made these disagreements meaningful was that the coding agent had also recorded its judgment calls in advance. The agent had found the ambiguity, made a resolution, and had no way to check whether that resolution was correct. The test suite was the check.
### Two Disagreements, Two Lessons
The two tasks disagreed in notably different ways, and both outcomes proved worth having:
– **Task one** revealed a genuine defect: the implementation used banker’s rounding (round half to even) in a financial context where half-up rounding was the correct business requirement. The acceptance criteria made the correct behavior explicit, and the test suite caught the mismatch immediately.
– **Task two** revealed no defect at all — but a decision that nobody had made, sitting unnoticed in the difference between `<` and `<=`. The implementation chose one boundary condition; the specification implied another. Neither was wrong on its own, but the mismatch surfaced a silent assumption that would have been impossible to detect without an independent test.These results demonstrate that test suites generated from specifications the code writer never saw *do* find meaningful discrepancies. Whether they find *more real defects* than conventionally written test suites remains an open question — and the one the author acknowledges has not yet been definitively answered.## Introducing qikly: A Tool for Independent Test GenerationTo explore these questions at scale, the author built **qikly**, an open-source Python library designed for spec-driven test automation with an explicit focus on separation of knowledge.The library's architecture enforces the principle that requirements and acceptance criteria are handled as distinct inputs. The coding agent receives only the requirements. The test-generation agent receives only the acceptance criteria. Neither has visibility into the other's domain.### A Practical Innovation: `--propose-fixtures`During experiments, the author identified a practical shortcoming: a criterion is only as strong as the data that exercises it. A test criterion about scientific notation cannot fire on datasets that contain no scientific notation, and a test written from that criterion will pass regardless of what the code does.This led to the addition of a new flag, `--propose-fixtures`, which analyzes the acceptance criteria and identifies gaps where the provided test data cannot satisfy a given criterion. The flag helps ensure that specifications are actually exercised, rather than silently passing by default.## Why the Distinction Matters Now### The Agent ProblemTraditional software development has always had safety nets. When a human writes both a specification and implementation, code review, peer testing, and diverse perspectives serve as implicit checks. Even when those checks fail, the human writing the test brings independent judgment and contextual awareness.An AI agent meets none of these checks. It does not have a colleague who reads the same sentence differently. It does not undergo code review in the traditional sense. If the same agent holds both the specification and the implementation in its working context, the "green test suite" may simply mean that the implementation matches its own assumptions — a tautology with no informational value.### What the Google Study Got RightThe Google study deserves credit for measuring something real and important. The finding that writing a specification down improves test quality is now supported by serious methodology: a real bug corpus, real statistics, and no stake in the researcher's particular conclusion. It validates the first half of the independence argument — that specification beats impression.But it stops at the boundary of what it can measure. The study does not address what happens when the specification and the code share a common origin. It does not test whether withholding the acceptance criteria from the coding agent changes outcomes. And it explicitly frames its contribution as improving the agent's reasoning, not as establishing an independent standard.### What Still Needs MeasurementThe open question remains: does independent test generation — where the test writer has no access to the implementation — catch more real-world defects than test suites written with full visibility into the code? This is the experiment that qikly is built to run, and it is the question that the research community has not yet answered definitively.## FAQ**What is spec-driven test generation?** Spec-driven test generation is a process where a formal specification of expected software behavior is used as input to automatically generate test cases. The specification describes what the software should do (pre-conditions, post-conditions, edge cases, and business rules), and an AI or automated tool translates that description into executable tests.**How does Independent Spec-Driven Development differ from traditional SDD?** In traditional SDD, the same specification is consumed by every agent in the pipeline — the planner, the test generator, and the code writer. In ISDD, the specification is deliberately split: the code-writing agent receives only the requirements (the decisions), while the test-writing agent receives only the acceptance criteria (the consequences). Neither has access to the other's portion.**Why can't a test suite that is generated from the code itself catch bugs in that same code?** If a test suite is generated by reading the implementation, it documents what the code does — including any bugs, edge-case mishandlings, or ambiguous decisions the developer made. The tests will pass on those behaviors by construction. A test can only fail if it asserts something *different* from what the code does, which requires an independent source of truth about what the code *should* do.**What is the `--propose-fixtures` flag in qikly?** The `--propose-fixtures` flag analyzes acceptance criteria and identifies which ones cannot be exercised by the available test data. For example, if a criterion references scientific notation but the provided data set contains no numbers in scientific notation, that criterion will never fire. The flag flags these gaps so users can supplement their test data.**Does the Google paper prove that ISDD catches more bugs?** No. The Google paper measures a related but distinct claim: that writing a specification down (rather than leaving it implicit) improves test-driven bug detection. It does not measure whether withholding the acceptance criteria from the code-writing agent yields better results. That question remains open.**Can ISDD be applied to non-AI development workflows?** The principles of ISDD apply whenever the source of truth for expected behavior is logically separated from the implementation process. In human teams, code review and diverse perspectives provide a degree of independence. ISDD makes that independence explicit and structural rather than relying on social and procedural checks.**What programming language is qikly written in?** qikly is a Python library, designed to integrate into Python-based development workflows and test pipelines.**Where can I find qikly?** qikly is available as an open-source project. The library is intended for developers and researchers exploring spec-driven test automation with a focus on separating requirements from acceptance criteria.## ConclusionThe evolution of software testing has moved steadily upstream — from executable tests, to shared natural-language scenarios, to machine-readable specifications that AI agents consume directly. Each step has improved clarity, communication, and automation. But a new challenge has emerged alongside AI-driven development: when the same agent both creates the specification and writes the code, the test suite loses its ability to serve as an independent judge.The research community has now measured that writing specifications down improves bug detection. What remains unmeasured — and what qikly is built to investigate — is whether *hiding* the specification from the code writer changes the outcome. Early experiments are promising: test suites generated from independent acceptance criteria consistently disagreed with implementations, surfacing both genuine defects and silent, unresolved decisions.Independent Spec-Driven Development is not yet proven to be superior in all contexts. But it addresses a real structural vulnerability in AI-assisted workflows — the conflation of what the code does with what the code should do. As AI agents take on larger roles in the software lifecycle, the question of *who knows what* becomes as important as the question of *who does what*.The next step is measurement at scale, across diverse codebases and real defect corpora. The tools and frameworks are ready. The open question is whether the research community will take it on.Thank you for reading



