Specs grounded in the codebase
Before an agent writes code, the ticket is turned into a spec that names the real files, patterns and conventions involved. Most agent failures start with a vague ticket.
Your developers have tried AI coding agents. They're impressive on a greenfield demo and unreliable on the ten-year-old system that pays the bills. The fix isn't a better model. It's a better harness.
Agentmodelharness
Everyone has access to the same models. What separates a team that ships with agents from one that cleans up after them is everything around the model: how work is specified, checked, proven and approved.
Before an agent writes code, the ticket is turned into a spec that names the real files, patterns and conventions involved. Most agent failures start with a vague ticket.
A separate reviewer with fresh context checks the work against the spec. The agent that wrote the code doesn't grade its own homework.
"Tests pass" comes with the test output. "This is handled" comes with the line that handles it. Reviewers check evidence, not assertions.
An engineer approves every merge. Agents add throughput; your team keeps the judgment and the accountability.
Your conventions, your architecture and your hard-won lessons, written down where agents can use them.
Every week, look at where agents went wrong and fix the harness, so the same mistake doesn't happen twice.
PRFlow is an open-source harness that makes AI coding agents dependable on real, mature codebases. It's an open-source project I lead, in daily use by dozens of developers.
You can read every part of it: the spec grounding, the fresh-context review, the evidence rules, the merge gate and the self-improvement loop. Use it as is, or as a reference for your own.
PRFlow on GitHubNo big-bang rollout. A pilot proves the approach on your code, with your team, before anyone scales it.
A real one, with real history. We map its conventions and set up the harness around it.
Not a toy. A ticket your team would normally pick up, run end to end through the harness.
One habit your team adopts and keeps, like grounded specs or the weekly improvement loop.
Six quick questions. You'll get a readiness band and what to fix first.
Thirty minutes with Daniel. Tell me about your codebase and how your team works today; I'll tell you honestly where agents will help and what to fix first.