live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
newsAI

Novita open-sources Chord kernels for INT4 mixture-of-experts serving in vLLM

Chord adds shape-aware W4A16 CUDA paths for Kimi K2.x, while its grouped vLLM integration remains unfinished.

Performance chart for Chord versus Humming on H200 and B300.
Chart: figures from the story
By The News Desk· Sep 17, 2026the quick take — two AI hosts go live when you do

Novita AI has open-sourced Chord, a set of CUDA operators for serving INT4 mixture-of-experts models with BF16 activations, and connected its indexed execution path to vLLM’s existing Humming backend. In a September 15 engineering report published by vLLM, the developers describe kernels tuned around the routed-token shapes seen when serving Kimi K2.x models.

Two paths for different serving shapes

Chord ships two kernel families. Its indexed path consumes vLLM’s sorted expert-routing buffers and covers H200 expert-parallel prefill, H200 tensor-parallel serving, H200 decode and B200/B300 decode. Separate grouped contiguous and grouped masked paths target prefill and decode on Hopper-class GPUs with different routing and weight layouts, according to the project.

The design responds to a practical problem in mixture-of-experts serving: prefill and decode send very different numbers of tokens to each expert. Chord chooses schedules from the routed shape it receives rather than treating total token count as the only useful signal. The implementation uses different tile sizes, occupancy limits and pipeline depths across those regimes.

What the measurements show

Against the matching public Humming paths, the project reports per-layer gains of 1.11 to 1.20 times for H200 expert-parallel prefill and 1.16 to 1.24 times for H200 expert-parallel decode. In an earlier end-to-end Kimi K2.6 serving run on eight H200 GPUs, the indexed path improved combined prefill throughput by 9.6% and decode output throughput by 4.1% to 8%, depending on batch size, the report says.

The largest published figure, 2.15 times for B300 decode, needs a qualification: Chord was compared with Humming’s default untuned configuration because the public Humming project did not ship a tuning table for that GPU. The authors also say the measurements are kernel-level results, not a guarantee of equivalent end-to-end gains for every deployment.

What operators can use now

The indexed path can be installed as a Python package and selected in compatible vLLM revisions with the existing humming quantization backend. It supports the compressed-tensors INT4 group-32 checkpoint format used by Kimi K2.x and rejects unsupported schemes instead of silently choosing an incompatible kernel.

The grouped operators are less mature. Their standalone API is available, but integration with vLLM’s Humming backend is still work in progress. Platform teams evaluating Chord should therefore distinguish the usable indexed route from the grouped roadmap, then benchmark complete serving behavior—including routing, communication and activation costs—on their own model and concurrency mix.

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.