KServe 0.20 gives model-serving teams safer rollouts and a hardware-backed path for encrypted models
The release adds TEE-based model decryption, progressive traffic controls, CPU-backed KV-cache offload and first-class tracing for Kubernetes inference workloads.
KServe 0.20 pulls several day-two model-serving controls into its Kubernetes APIs: encrypted-model handling inside trusted execution environments, progressive traffic management, CPU-backed KV-cache offload and declarative tracing. The project’s release post was published August 6.
What changed
For sensitive model artifacts, KServe can now download JWE-encrypted files and ask a local Confidential Data Hub inside a trusted guest to obtain the decryption key through hardware attestation. The model is decrypted inside the trusted execution environment rather than exposed to the host. The implementation works with both InferenceService and LLMInferenceService, and the project says it is compatible with Key Broker Service backends including Trustee and Intel Trust Authority.
The release also adds weighted traffic groups to LLMInferenceService. Operators can shift requests between service versions while readiness and degradation are tracked for each group independently. For conventional InferenceService workloads in RawDeployment mode, KServe 0.20 adds named canaries with percentage-based traffic allocation; promotion can retarget the stable service without recreating model pods.
For vLLM deployments, a structured kvCacheOffloading.cpu field now lets operators reserve CPU memory as a secondary KV-cache tier. KServe renders the corresponding vLLM transfer configuration, including for disaggregated prefill/decode layouts. A separate tracing API enables OpenTelemetry for the inference server and scheduler without requiring teams to inject environment variables through pod-template overrides.
Who should care
Platform teams already standardizing inference behind Kubernetes custom resources get the clearest benefit. The release moves rollout safety, model confidentiality and observability into declarative configuration instead of leaving each model team to assemble sidecars, environment variables and bespoke deployment logic.
There are useful runtime changes too. KServe now supports vLLM as a standalone InferenceService runtime, routes Anthropic Messages API traffic through the inference pool, and can mount OCI model images directly with Kubernetes ImageVolume. That native OCI path requires Kubernetes 1.33 or newer.
What to do
Teams testing 0.20 should start with one operational problem rather than enabling everything at once. Canary or weighted traffic controls are the lowest-risk place to validate rollback behavior. KV-cache offload should be benchmarked against latency and host-memory pressure under realistic prompts. Confidential serving needs a working attestation and key-broker chain, so it belongs in a threat-model review rather than a routine configuration toggle.
The release also upgrades its llm-d dependency to 0.8.0 and adopts llm-d.ai routing CRDs. Existing users should therefore review CRD and controller changes before upgrading production clusters.
sources
- Announcing KServe v0.20kserve.github.io
comments · 0