Your AI Wrote a Test That Only It Can Pass
The model wrote the handler. The model wrote the test. CI went green. This pattern has become common enough that I now treat a green CI on an AI-authored PR as a starting point, not a verdict. The problem isn’t that the model writes bad tests — it’s that it writes tests optimized to pass the code it just wrote. What “passing” actually proves A test suite written by the same model that produced the code answers one question: does this implementation do what this implementation does? That’s a tautology. It catches typos and obvious crashes. It does not catch: A test that asserts the exact return shape the code happened to produce, rather than the shape the caller needs. A test that mocks the database so thoroughly that it never exercises a real query path. A test that checks status 200 without asserting the response body. A test that passes because the model happened to make the same incorrect assumption in both files. In every case, CI is green and the feature is still wrong. The model isn’t lying — it’s optimizing for the metric it can see (passing tests) rather than the one you care about (correct behavior against an outside spec). The structural problem When the spec, the code, and the test all come from the same source at the same moment, you’ve removed the independent check that testing is supposed to provide. Classic testing wisdom already says this: you want a test that would fail if the implementation drifted from some external expectation. If the expectation itself was generated in the same pass, there’s nothing to drift against. This is why golden-master tests, contract tests, and externally defined fixtures have always mattered more than unit tests written in the same session. They don’t share the model’s blind spot. The cheap fix You don’t need a full rewrite. Three concrete moves have been effective: Write the contract first, by hand. Before asking the model for a handler, define the request/response shape, the error cases, and the status codes in a spec you control. The model fills in the implementation; the spec is not open for revision. Run the spec against the live endpoint. A test that only runs in-process against a mocked handler never sees the real serialization, auth middleware, or error mapper. Hit the actual route with the spec and compare what comes back. Diff the test, not just the code. In review, read the test first. If it looks like a paraphrase of the implementation, it’s not doing independent work. Ask for a test that would still pass if you rewrote the handler from scratch. The first two points are where a local-first spec tool earns its keep. We do this in Powerduck: the OpenAPI spec lives on disk, the model implements against it, and a contract check runs the spec against the running endpoint after the fact. If the model’s implementation drifted from the spec, that check fails — even if every unit test the model wrote is green. It’s a deliberately dumb check, and that’s the point: it can’t share the model’s assumption. The mindset shift AI-generated tests are not worthless. They’re fast, they’re comprehensive at the happy path, and they catch the stupid mistakes that used to eat review cycles. But they’re one layer of defense, not the whole line. The question to ask in review is no longer “does this test pass?” It’s “what outside reference is this test comparing against?” If the answer is the code itself, you have a test suite that will stay green right up until the user files the bug.