live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
releaseAI

vllm-metal turns Apple Silicon into a concurrent local serving target

The vLLM plugin adds paged KV caching, packed attention and an OpenAI-compatible server for multi-request inference on Macs.

Old padded batching versus vllm-metal packed paged inference on Mac.
Side by side: what changed
By The News Desk· Sep 24, 2026the quick take — two AI hosts go live when you do

vLLM has made Apple Silicon a first-class target for concurrent local inference through vllm-metal, a plugin that combines vLLM’s serving machinery with Apple’s MLX and Metal execution stack. The project’s September 22 engineering post describes v0.28.0 as its first official release; the current Homebrew installation path provides v0.29.0.

What changed

The plugin reuses vLLM’s V1 scheduler, paged key-value cache management, chunked prefill, sampling and OpenAI-compatible frontend. Model layers come from mlx_lm, while vllm-metal replaces attention with a paged variable-length Metal kernel that preserves request boundaries inside a packed batch.

That division matters for concurrent workloads. Rather than padding every prompt to the longest request in a batch, vllm-metal concatenates scheduled query tokens and uses sequence offsets to keep requests separate. Its fixed-size KV pages let admitted requests grow without repeatedly reshaping a contiguous cache. A memory guard also reserves headroom for macOS and other applications, then sizes a fixed KV cache after a startup warmup.

The first official release added batched multi-token prediction, GGUF checkpoints, hybrid-attention model support and an accelerated prefill path for M5 systems. The documented feature set also includes LoRA adapters, structured outputs, pipeline parallelism across Macs, and experimental vision-language, embedding, reranking and speech-to-text paths.

Who should care

This is primarily a developer and workstation-serving release, not a replacement for clustered production inference. It gives teams using Macs an OpenAI-compatible endpoint for local coding agents, experiments and multi-user test workloads while retaining vLLM’s admission control and batching model.

The project reports lower time to first token and end-to-end latency than the compared engines at concurrency levels two and four for a 4-bit Qwen3.8-27B workload on an M5 Pro with 64 GB of memory. Those are project-run benchmark results rather than an independent evaluation, and the post notes that results cover completed requests using engine-specific 4-bit conversions.

What to try

On macOS 15 or later, the documented path is to add the project’s Homebrew tap, install vllm-metal, and launch a model with vllm serve. Existing clients can then use the local OpenAI-compatible /v1 endpoint.

Teams should start with an explicit --gpu-memory-utilization budget and reproduce the project’s benchmark against their own prompt lengths and concurrency. Several paths remain bounded: batched MTP is currently limited to Gemma 4 with greedy sampling and synchronous scheduling, while prefix reuse for Qwen hybrid models is experimental and cannot yet be combined with speculative decoding.

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.