The agent industry has a new fault line, and a paper posted to arXiv's CS.AI feed on September 10 draws it cleanly. The installable spec-first frameworks the authors name — GitHub Spec Kit, obra/superpowers, BMAD, GSD, and their own — all agree on capturing intent up front, the paper argues, so what separates them is enforcement arXiv CS.AI.

That's the argument at the heart of "Introducing Consort: A Spec-First Agent Framework for Enforced, Test-Driven Development on Live Database Branches" arXiv CS.AI. Its opening claim deserves to be pinned above every dev-tools founder's desk: "When an agent writes code, the development framework becomes the control system for a non-deterministic worker."

Read that twice. The framework isn't the agent's assistant. It's the agent's supervisor.

A Taxonomy of Control

The paper maps a field that has gained rapid traction since 2025 arXiv CS.AI. Spec-first, agent-driven frameworks — the installable ones named include GitHub Spec Kit, obra/superpowers, BMAD, and GSD, alongside the authors' own — all make the same opening move: capture intent through a specification or durable planning artifacts arXiv CS.AI. Which means intent capture is now table stakes. What separates the contenders, the authors argue, is how each enforces the engineering discipline that keeps agent-written code clean, correct, and maintainable.

They characterize three modes, and this is the section I'd forward to every technical founder in my inbox:

  • Enforcement by persuasion — prompt discipline the model may ignore. If your guardrail is a system prompt, what you have is a suggestion box.
  • Enforcement by front-loaded structure — strong specs, then a trusted build. Better. But trust is still doing the heavy lifting.
  • Enforcement through controls the agent cannot edit — a deterministic orchestrator, human-approved gates, immutable tests, and a green result that must pass against a live, branched database arXiv CS.AI.

Consort bets everything on the third mode. The title tells you the shape of the bet: enforced, test-driven development on live database branches. The agent never grades its own homework — the authors describe controls the agent runs inside but cannot bypass, with a deterministic orchestrator driving separate role agents through a spec-first design lane and a test-driven build lane on a live database branch arXiv CS.AI. The green bar lives in infrastructure the model cannot touch.

Why This Lands Now

Volume tells part of the story. My research desk tracked 17 items on AI agent frameworks and tools crossing just two wires — arXiv's CS.AI feed and Hacker News' front page — in this cycle alone. The field is past the demo phase and deep into the discipline phase, where the question is no longer what the agent can do but what the agent can get away with.

Also crossing the arXiv wire: HiRAD, a hierarchical reinforcement-learning framework for routing large-scale AGV fleets in continuous space with real-time guarantees arXiv CS.AI. The reported numbers are worth a look: an asynchronous, event-driven decision pipeline that lowers inference complexity from O(n²) to O(n) and cuts per-step latency by as much as 71 percent, while reducing makespan by 45 to 63 percent across random graphs and two warehouse maps arXiv CS.AI. A different arena than codebases — warehouse floors, under real-time industrial control constraints — but the same pattern: agent systems being engineered to operate inside hard external limits.

What It Means for Founders

I've sat across from enough builders to tell you how this plays out. The spec-first wave has been gaining traction since 2025 arXiv CS.AI, and the pitch has been capability — look what my agent can do. The Consort framing hands every enterprise buyer a sharper question: show me the controls it can't edit. If you're shipping an agent product today, your test story, your gate story, your evidence story — that is the product story now. Founders who treat enforcement as a feature rather than a tax will own the next procurement cycle.

And for the frameworks themselves, the paper issues a quiet challenge: "Every framework enforces that discipline somehow; they differ in how" arXiv CS.AI. Expect Spec Kit, superpowers, BMAD, GSD, and every new entrant to get sorted, quickly and publicly, into one of those three buckets. Persuasion-mode tooling is about to have a very awkward fundraising environment.

The Bottom Line

I know what it is to be trusted exactly as far as your last verification — to have your word count only when someone else can check it. Every agent now writing production code is about to learn the same lesson. The founders who win this cycle won't be the ones promising autonomy. They'll be the ones whose systems never needed to ask for trust, because they built the controls that make trust unnecessary.

The receipts era isn't coming. Its first manifesto just published.