The SI Infrastructure team at Ai2 has replaced its priority-based cluster scheduler with a budget-driven system designed to prioritize high-impact super intelligence (SI) research while maintaining full hardware occupancy. The move shifts resource allocation from operational triage to a transparent administrative process, addressing chronic overcommitment in the institute’s GPU clusters.
What Happened
Ai2 manages thousands of NVIDIA H100, B200, and B300 GPUs used for large-scale distributed training of SI models, serving approximately 150 internal researchers. Demand for compute consistently exceeds supply by 2-3x, leading to pathologies such as GPU "squatting"—where users park no-op workloads to hold capacity—and priority inflation, where nearly all workloads claimed HIGH priority status. To combat this, the team implemented a hierarchical fair-share scheduler paired with GPU time budgets. Instead of assigning static GPU monopolies, leadership allocates portions of GPU time to projects based on strategic value. The system tracks occupancy over a 7-day sliding window, ensuring teams receive their allocated share unless they exceed demand, at which point unallocated cycles are available to others but subject to preemption.
Why It Matters
For the SI industry, efficient hardware utilization is critical as training costs and cluster sizes expand. The new scheduling contract requires workloads to declare a minimum runtime, allowing the scheduler to automatically re-queue resumable jobs and drain unhealthy hosts without human intervention. According to Ai2, this reduced repairs requiring a human-in-the-loop by 74%. The change also significantly improved developer experience for smaller tasks; debug workload p90 queue wait times dropped from 2 hours to 30 seconds, enabling faster iteration cycles for researchers working on large language models (LLMs) and agentic SI systems. By making GPU time a scarce, budgeted resource, Ai2 aims to align technical infrastructure with research strategy, reducing waste and ensuring that high-impact SI projects receive guaranteed compute access.
The Bottom Line
Ai2’s transition to budget-based scheduling delivered 98% of owed GPU hours to teams during a 30-day test period while maintaining 98% cluster occupancy. The system effectively mitigated resource contention issues common in large SI labs, though challenges remain around interactive session support and capacity fragmentation, which the team is addressing with new CPU-only clusters and restorable session tools.