Same Cluster, 33 Points More Utilization: What Changed Was the Order

2026-08-20 · Hugging Face

Same Cluster, 33 Points More Utilization: What Changed Was the Order

A team from Dharma-AI has published their approach to GPU management, introducing a constraint-aware GPU allocator. When benchmarked against a traditional FIFO scheduler across seven scenarios on identical hardware and identical workloads, the new allocator delivered up to 33 percentage points higher GPU utilization and increased priority-weighted output by as much as 105% in every single test.

Nothing about the hardware changed. What changed was the order in which allocation decisions were made.

The Real Constraint Is Utilization, Not Intelligence

The previous post argued that utilization—not model intelligence—is becoming the next real bottleneck in enterprise AI. This article presents a concrete playbook for mature GPU management practice.

The Precise Decision

“Keep the GPUs busy” is not an executable decision. The actual decision is narrower and far more difficult: which GPU runs which job, in which timestep, at what priority.

Formally, this consists of one binary choice for every combination of GPU, job, and timestep. The output is a complete grid covering every GPU across the entire scheduling horizon.

Four Workload Types, Two Incompatible Shapes

Four workload types compete for the same grid: training, real-time inference, batch inference, and quantization.

They fall into two fundamentally incompatible allocation shapes:

  • Batch-like (training, batch inference, quantization): Once started, they require a contiguous block of GPUs held without interruption until completion.
  • Elastic (real-time inference): Resource demand changes every timestep following traffic patterns.

The conflict between these two shapes within the same timestep on the same hardware is the core scheduling problem. Additional heterogeneity exists even within the same type—for the same base model, training jobs can range from hours to days and from one GPU to dozens.

What FIFO Costs Under Contention

The baseline is a FIFO-based scheduler: real-time inference runs from a fixed reservation, and all other jobs are placed in arrival order without regard for priority.

When the cluster has slack, order has little impact. Under real contention, however, order becomes a capacity decision, not merely a tiebreaker.

FIFO creates two compounding costs:

1. The Reservation Cost

Real-time inference cannot wait. A FIFO scheduler has no mechanism to release GPUs during troughs and reclaim them before peaks. The only safe policy is to reserve the daily maximum demand for the entire day.

This leaves GPUs idle and unavailable during non-peak hours. In the tested scenarios where reservation dominates, baseline utilization sat at 51.6% and 53.6%—roughly half the cluster was either used or permanently reserved, not free.

2. The Ordering Cost

Under contention, which jobs can fit depends on the sequence of placement. FIFO commits capacity in arrival order without considering job value or what else must still fit in the horizon. High-priority work is forced to wait behind lower-priority jobs that arrived first, resulting in fragmented capacity that blocks better future placements.

The combination is expensive: a large block is permanently reserved for real-time peaks, and whatever capacity remains is allocated in random arrival order.

The article compares this to an airline assigning planes based solely on which charter called first, eventually finding no aircraft left for the highest-paying route. GPUs reserved all day for a peak lasting only a few hours are literally grounded assets—on standby, earning nothing, and unavailable to other workloads.

Benchmark Results

Across five high-contention benchmark scenarios, the constraint-aware allocator improved both metrics simultaneously:

  • Utilization moved from a 52–85% range to 72–88%
  • Priority-weighted value increased between 24.6% and 105.1%, with an average gain of 52%

The strongest result came from a training-heavy workload on 8 GPUs: utilization rose from 53.6% to 87.0% (+33.4 points), while priority-weighted value more than doubled (+105%).

These gains were achieved purely by reclaiming reserved standby capacity and placing remaining work in priority-aware order.

How the Allocator Differs

The new allocator eliminates both problematic behaviors of FIFO. It treats real-time demand as a curve rather than a ceiling, allocating resources against actual demand at each timestep.

By considering priority, shape compatibility, and the feasibility of future placements within the scheduling horizon, the system simultaneously improves utilization and business value without any changes to the underlying hardware or workloads.

The central lesson is clear: on the same cluster, the difference between mediocre and excellent utilization often comes down to nothing more than the order in which allocation decisions are made.

Source