On this page

Blog

Hand the agent one build spec and count the questions.

We put one of our own delivered build specs in front of a coding agent cold, told it to ask rather than build, and counted. Twenty-nine to thirty-nine questions a run, about fifteen of them real gaps. The count is a better readiness test than any review score, and it costs one session.

Last week we put one of our own build specs in front of a coding agent, cold, and told it to ask rather than build. Fresh session, no access to the repository, nothing but the spec. Three runs. It came back with twenty-nine to thirty-nine questions each time.

The spec wasn't a napkin. It was a delivered user story from an internal project: three acceptance criteria, the written context for the component it belongs to, and the three architecture decisions it leans on. Working code had already been built and delivered against it. By any review we'd have given it, it read fine.

Not every question counts, so we sorted them. Real gaps, where the spec does not answer the question. Preferences, where the spec deliberately leaves the call open. Noise, where the answer was already in the packet. About fifteen a run were real gaps. Across the three runs those came to nineteen distinct things the spec never said.

That number is the readiness test. A review score tells you the spec reads well. A question tells you exactly where it stops. It's a boring number, which is what I like about it. You can run it on a real ticket this week, and the result is a list, not an opinion.

Why the agent is the right reader

A colleague reviewing the same spec would have asked almost none of those questions, which is exactly why a colleague is the wrong reader for this test. They know the directory layout. They know which error format the team uses. They've seen the manifest file. Every gap in the document gets filled from the corridor before anyone notices it, and the spec gets waved through.

An agent in a fresh session has no corridor. Whatever the spec doesn't say, it doesn't know, so its questions are the spec's gaps and nobody else's. The condition has to be real, though. Give it the repository and it stops asking, because it reads the code instead, and you've measured the codebase rather than the document. Cold means cold: one spec, one session, ask rather than build.

A cold reader's questions are the spec's gaps Two panels, one build spec. The spec is a document with three blank slots between its lines of text: the things it never says. On the left, a colleague reads the spec. A box named Corridor, the team's knowledge, sends an arrow into each blank; the blanks fill, and an arrow leads from the spec to a tick: waved through. On the right, an agent reads the spec cold, inside a box named Fresh session with nothing else in it and no corridor. The three blanks stay open, drawn in blue, and an arrow leads from each blank out of the session to a question card. The agent's questions are the spec's gaps, one for one. The drawing shows the mechanism, not the count. A cold reader's questions are the spec's gaps A colleague reads the spec Corridor Teamknowledge Build spec Wavedthrough An agent reads the spec cold Fresh session Build spec ? ? ? Questions

The gaps clustered where you'd fear. How the schema library exposes its schemas, and how parent and child records point at each other. Where things live on disk, and where the root path comes from. Which document kinds exist, and which error categories have to be told apart. What happens when a tree walk hits a broken reference, whether that's abort, partial result or skip. Whether a write is atomic, and whether a reference outside the tree gets followed or refused. These are the places a building agent makes a call without cover, and where a silent guess is expensive to unwind later.

Fix the spec, not the chat

Every question is an edit to the document, never an answer typed into the conversation. An answer in the chat lives for one session. The next session, or the next colleague, meets the same gap and gets a different answer, or guesses. Put it in the spec, run it cold again, count again.

To be honest I expected ours to do better than that. This one already had working code delivered against it. The agent still found about fifteen things a run the document never said, and the code built against it had to settle every one of them. Those calls were made. They just weren't written down where the next reader could find them.

Which is why the count has become a habit for me rather than a one-off. The document is the source of truth and the conversation is scrap paper; anything that only exists in a chat might as well not exist. That's the RCF methodology, the way I build software with agents, in one sentence, and the written build spec is the piece of it this test checks. Under that method the document is called a functional build spec. You don't need the method or the name to run the test.

What zero means

Zero means complete enough to start. It does not mean right. The count checks whether the document answers the questions a builder has to ask; it says nothing about whether the requirement behind it was worth building. A spec can be complete and wrong, and the agent will build the wrong thing without a single question.

A confident agent can also just not ask. A model told to build will happily assume and carry on, which is why the instruction is ask first, and why one clean run isn't proof. Run it more than once. Our three runs overlapped heavily but not completely; the most conservative found fourteen real gaps and the most thorough sixteen.

Two honest limits on our numbers. One model family, one spec, so I'm not claiming the counts generalise. And no control run yet. I think a spec that carries its failure cases, its closed sets and its must-nots as first-class fields drops the count hard, and I've started writing them that way, but I haven't measured it. That's the next test, on the same story, once one exists in that form.

One ticket, a fresh session, only the spec. Tell the agent to ask, not build. Count the real gaps. Fix the document, not the chat. Run it again. And the next time someone tells you the agent got it wrong, ask what the spec said first. In my experience the answer is often that it didn't.

Barry