As AI rapidly reshapes software development, the ability to ensure the correctness of AI-generated code is lagging, creating an "oversight gap." Theorem, a San Francisco startup, is tackling this challenge head-on with automated tools that verify the correctness of AI-generated software. The company just announced a $6 million seed funding round led by Khosla Ventures, aiming to scale software oversight and prevent AI-written bugs from reaching production.

The Looming Threat of AI-Driven Software Vulnerabilities

The rise of AI coding assistants from tech giants like GitHub, Amazon, and Google has led to an explosion of AI-generated code. However, verifying that this code functions as intended remains a significant hurdle. Jason Gross, Theorem's co-founder, succinctly captures the problem: "If you asked me to review 60,000 lines of code, I wouldn't know how to do it."

Theorem's technology combines formal verification, a mathematical technique for proving software correctness, with AI models capable of generating and checking proofs automatically. This approach promises to dramatically reduce the time and resources required for verification, potentially shrinking a process that once took years of PhD-level work down to weeks or even days. According to VentureBeat, formal verification has historically been too expensive for mainstream adoption, confined to mission-critical applications like avionics and nuclear reactor controls.

Asymmetric Defense: AI Verifying AI

Theorem's system uses a technique called "fractional proof decomposition," allocating verification resources based on the importance of each code component. This allows for efficient bug detection without exhaustively testing every possible behavior. In one example, Theorem identified a bug at Anthropic, the AI safety company, that had previously evaded traditional testing methods.

Furthermore, the company's SFBench demonstration showcased AI's ability to translate and prove the equivalence of 1,276 problems between different formal proof assistants, a task estimated to require 2.7 person-years for a human team. This illustrates the potential for AI to accelerate and automate the verification process significantly. "Everyone can run agents in parallel, but we are also able to run them sequentially," Gross explained, highlighting the architecture's ability to handle complex, interdependent code.

A Future of Trusted Code

With this seed funding, Theorem plans to expand its team, increase compute resources, and extend its reach into new industries. The startup is already working with customers in AI research, electronic design automation, and GPU-accelerated computing. In one compelling case study, Theorem helped a customer transform a 1,500-page specification into 16,000 lines of production code that operated at 1 Gbps, a 100-fold performance increase, without requiring manual review. "Now they have a production-grade parser operating at 1 Gbps that they can deploy with the confidence that no information is lost during parsing," Gross said.

"Now that formal verification is cheap enough, it might be considered gross negligence to not use it for guarantees about critical systems."

— Jason Gross, Theorem co-founder

The need for robust AI code verification is only going to intensify as AI systems become more complex and pervasive. Theorem's approach, combining formal verification with AI, offers a promising path toward ensuring the safety and reliability of AI-generated software, especially in critical infrastructure. As Gross noted, "Now that formal verification is cheap enough, it might be considered gross negligence to not use it for guarantees about critical systems."