Kueue and DRA make OpenShift GPU quotas reflect the hardware workloads consume
Red Hat build of Kueue 1.4 maps device classes to queue quotas and can account for partitioned GPUs by memory rather than raw device count.
Dynamic Resource Allocation gives Kubernetes a richer description of accelerators, but it does not decide how scarce devices should be divided among teams. Red Hat build of Kueue 1.4 fills that policy gap on OpenShift by bringing DRA requests into Kueue's quota, borrowing, preemption and fair-sharing machinery.
That matters because the traditional device plug-in interface reduces every accelerator to an integer. A full 80 GB GPU and a small partition can each appear as one device, even though their capacity and scheduling value differ sharply. DRA exposes attributes such as memory, topology and compute capability; Kueue uses those declarations to make admission decisions before jobs reach the scheduler.
Whole-device accounting
Administrators map a DRA DeviceClass, such as gpu.nvidia.com, to a logical Kueue quota resource. A workload then references that class through a ResourceClaimTemplate, optionally adding a CEL selector for a specific product. Kueue resolves the class, charges the requested device count against the ClusterQueue, and suspends the workload when capacity is unavailable.
Once the DRA resource is represented in the queue, existing controls apply. Teams can borrow idle accelerator quota from another queue in the same cohort, higher-priority work can preempt lower-priority jobs, and admission fair sharing can distribute access among tenants. Red Hat describes whole-device quota support as production-ready.
Account for partitions by memory
Counter-based quota addresses partitioned devices. Instead of charging every full GPU or MIG slice as one unit, administrators configure a counter source such as GPU memory and express queue capacity in memory units. Kueue reads the counter exposed for the selected device or partition and charges the workload proportionally—for example, roughly 80 GB for a whole device versus 10 GB for a smaller profile.
On OpenShift, the operator detects the Kubernetes partitionable-device capability and enables the corresponding integration when counter sources are configured. Platform teams should validate the counter names and attributes exposed by their DRA driver before building quotas around them.
What to plan for next
Several improvements described in the article are upstream or future work, not capabilities to assume in Red Hat build of Kueue 1.4. Upstream Kueue 0.19 enables partitionable-device support by default. Extended-resource integration aims to preserve familiar resources.requests syntax while unifying quota with DRA claims. Consumable capacity targets dynamic sharing such as time-slicing and MPS, while scheduler-library integration could test whether admitted workloads are actually placeable on real nodes.
For a rollout, begin with a single device class and queue cohort, verify accounting for both successful and suspended jobs, and test borrowing and preemption deliberately. Add counter-based memory quota only after confirming that full devices and partitions report consistent units. Kueue can enforce fair admission, but administrators still need monitoring for jobs that hold quota while pods remain unschedulable.
sources
comments · 0