live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
newsSECURITY

One negative token ID can stop an exposed vLLM embedding server

The unauthenticated availability flaw reaches the embeddings and pooling APIs; vLLM 0.28.0 rejects the input before it reaches CUDA.

Before-and-after embedding server path with negative token rejected upstream.
AI-generated illustration
By The News Desk· Sep 13, 2026the quick take — two AI hosts go live when you do

A single unauthenticated request can terminate a vulnerable vLLM process when an embedding or pooling client supplies a negative token ID. The project’s security advisory says -1 is sufficient: the request receives an HTTP 500 response, the health endpoint then refuses connections, and the process must be restarted.

What fails

The affected /v1/embeddings and /pooling paths accept token IDs directly. vLLM checked whether an ID exceeded the vocabulary’s upper bound, but did not reject values below zero. A negative value therefore reached the GPU embedding-table lookup and triggered a CUDA device-side assertion.

That failure poisons the CUDA context rather than producing a recoverable Python indexing exception. According to the advisory, subsequent CUDA work in the same process fails, including requests unrelated to the malicious input. The impact is availability only; the project reports no confidentiality or integrity loss.

The issue is especially sharp at the service boundary because it requires no authentication, unusual payload size or request sequence. Restarting the server under a supervisor is not a sufficient defense: the same request can be repeated as soon as the process returns.

Who needs to act

Operators should treat internet-facing or otherwise untrusted embedding and pooling endpoints as exposed. The project verified the failure against an unpatched vLLM 0.27.1 image and verified the corrected behavior in 0.28.0.

vLLM 0.28.0 adds a lower-bound check before token IDs reach the device. Negative IDs now receive an HTTP 400 response, while the engine and health endpoint remain available. The validation path is shared across generation, embedding and pooling requests.

What to do

Upgrade affected deployments to vLLM 0.28.0 or later. Where an immediate upgrade is not possible, the advisory recommends rejecting negative integers in input at a reverse proxy or API gateway, accepting only string input from untrusted callers, or restricting network access to the two affected endpoints.

The ordinary completions endpoint is not affected because its request schema already constrains token IDs to non-negative values. The score and rerank endpoints do not accept token-ID input in this form.

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.