live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
newsAI

vLLM puts distortion-free text watermarking into the serving path

The new Gumbel-max implementation embeds a detectable provenance signal during generation while preserving sampling behavior and speculative-decoding performance.

vLLM watermarking in serving, not post-processing
AI-generated diagram
By The News Desk· Sep 25, 2026the quick take — two AI hosts go live when you do

vLLM now supports keyed text watermarking directly in its serving pipeline, adding a detectable provenance signal without deliberately biasing the model toward particular words or styles. The implementation uses Gumbel-max sampling, fused GPU kernels and special handling for speculative decoding, according to a technical account published by vLLM and co-authored by engineers from Mistral and Red Hat.

What changed

The design replaces ordinary random draws at each decoding step with reproducible keyed values derived from the recent token context and each candidate token. A detector that knows the secret key and tokenizer can reconstruct those values from generated text and calculate whether the observed sequence carries the watermark; it does not need access to the model weights or logits. In expectation over keys, the Gumbel-max procedure preserves the model’s original categorical sampling distribution, the authors explain.

vLLM integrates the feature into Model Runner v2 rather than adding a post-processing layer. Its GPU implementation fuses pseudorandom-number generation, the Gumbel transformation and the final argmax reduction, avoiding a temporary tensor spanning the batch and full vocabulary. In the published Qwen3.5-27B tests, matched throughput changes across batch sizes ranged from a 1.1% decrease to a 2.0% increase, with no consistent slowdown.

Why speculative decoding matters

Watermarking both the draft and target distributions with one key can reduce their overlap and cut the acceptance rate that makes speculative decoding useful. vLLM instead uses separate keys for accepted draft tokens and target residual or bonus tokens. Detection combines evidence from both keys, trading some signal strength for preserved acceptance behavior.

The implementation also skips watermarking when a recent context repeats. Without that safeguard, the same context and key reproduce the same random values, which can correlate later token choices and reduce sequence diversity. vLLM reports that this context-deduplication check kept the worst measured end-to-end throughput impact to 0.19% in its Qwen3.5-27B evaluation.

Who should care

Platform teams running vLLM can now test provenance marking at inference time rather than routing output through a separate service. The server accepts a watermark configuration containing the Gumbel algorithm and a key, while vLLM also provides a minimal HTTP detector example using the matching tokenizer and key.

The limits matter. Detection becomes stronger as more independent token evidence accumulates, so short or highly predictable outputs carry less signal. Operators also need key-management and tokenizer-matching practices, and testing against many candidate keys or configurations requires stricter statistical thresholds to control false positives. This is therefore an implementable serving primitive, not a universal declaration that any text came from an AI system.

sources

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.