The Allen Institute for AI (Ai2) disclosed on October 9 that it replaced the priority-based scheduler governing its thousands of GPUs with a budget-and-fair-share system designed to steer scarce compute toward the most valuable research.
The change addresses a persistent problem in academic and industrial AI labs: how to allocate over-subscribed GPU clusters when demand routinely outpaces capacity by two or three times, while preventing users from gaming the system.
Ai2’s infrastructure team manages clusters of Nvidia H100, B200, and B300 GPUs ranging from 88 to 1,024 accelerators, serving roughly 150 internal researchers, according to a blog post on Hugging Face. At any moment, outstanding GPU requests total 2–3 times the available capacity, the team wrote.
From Priority to Budgets
Under the old scheme, workloads were scheduled by priority and could opt out of preemption. That created “predictable pathologies,” the post said: researchers squatted on GPUs with no-op workloads so they could quickly connect when needed, and priority inflation eventually saw all jobs marked HIGH, starving lower priority levels. On-call engineers spent most of their ticket time negotiating shutdowns of non-preemptable workloads on hosts with known maintenance problems.
The team described the situation as a “tragedy of the commons” in which individuals competed for a shared resource and wound up degrading it for everyone. Assigning teams dedicated GPU monopolies solved some issues but caused idle cycles because research demand is bursty, leaving GPUs empty while other teams waited.
To break the cycle, Ai2 moved from handing out GPUs to allocating portions of GPU time. Instead of a static schedule, managers now decide in advance how to fund each research effort with a time budget, much like investors. The scheduler then uses those budgets to prioritize arriving workloads.
“We enabled leadership to think like investors,” the team wrote. “Before the workloads exist, decide how to fund each research effort with GPU time based on their judgment of its likely impact.”
How the Scheduler Works
The new system pairs a hierarchical fair-share algorithm with the budget allocations. The algorithm is similar to the Fair Tree method used in the Slurm workload manager, the post noted, and sorts workloads from under-utilized allocations above those from over-utilized allocations over a sliding seven-day lookback window.
Every GPU request must be funded by a budget to be protected from preemption; otherwise, it runs as unallocated work that can be evicted at any time. “Now, nothing is free, so any trick to get GPU time draws from the benefiting user’s allocation,” the team said. “A squatting workload is spending team budget on nothing.”
A “scheduling contract” requires each job to declare a minimum runtime during which it is protected. After that window, the scheduler can rebalance by preempting and re-queuing resumable workloads. The approach introduces time-slicing, preventing long-running jobs from blocking others and automating drain operations when a host needs maintenance.
Operational Results
The contract feature cut repairs requiring a human-in-the-loop by 74%, a “massive savings in on-call toil,” according to the post. Ai2 researcher Chris Clark said the new scheduler made it feel like the team had “an extra 30% compute” because bursty workloads could now reclaim unused budget time without delay or preemption.
The institute acknowledged that scheduling policy changes carry risks and ran simulations to test the new logic before deployment. No independent benchmark results were included in the blog post.
Budget decisions are made by the managers with the most context: lead researchers decide within a project, principal investigators across a program, and program managers or the CEO between programs. “There are frequent opportunities for researchers to advocate for the time they need,” the post said.