live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
analysisPLATFORM

Red Hat’s Krkn adds resiliency scoring and deeper chaos telemetry

The chaos-engineering project now measures weighted SLO failures, endpoint downtime and KubeVirt VM connectivity instead of relying on a binary pass/fail result.

Resiliency score and downtime telemetry across chaos tests.
Chart: figures from the story
By The News Desk· Sep 30, 2026the quick take — two AI hosts go live when you do

Red Hat has detailed a broader telemetry model for Krkn chaos engineering, combining a beta Resiliency Score with continuous HTTP health checks and KubeVirt virtual-machine checks. The change is meant to show how severely a system degrades during a fault experiment, not merely whether it remains nominally available.

A weighted score for the chaos window

Krkn’s beta Resiliency Score ranges from 0% to 100%. It evaluates Prometheus data collected during the chaos window against configured service-level objectives. A triggered alert marks the associated SLO as failed, and a weighted model determines the overall result.

The default weighting assigns one point to warnings and three points to critical outages, while teams can set custom weights and add SLOs. That gives operators a way to make failures in business-critical services count more heavily than lower-impact alerts. Krkn emits both an overall score and per-scenario reports showing passed and failed SLOs and points lost.

Endpoint and VM checks expose the duration

Continuous HTTP health checks probe configured application endpoints at user-defined intervals, with support for bearer tokens and credential tuples. Their telemetry records status codes plus the start, end and duration of downtime. Teams can therefore measure recovery time instead of stopping at a failed-health-check flag.

For virtualized workloads, Krkn’s KubeVirt checks monitor virtual-machine instances throughout the run. They test SSH connectivity through virtctl, or through worker nodes in disconnected environments, and record loss of connectivity, node placement, IP changes after rescheduling and downtime duration. A final post-experiment check flags VMs that remain unreachable after fault injection ends.

A CI quality gate

Red Hat positions the three layers—weighted SLO scoring, endpoint availability and VM connectivity—as a versioned quality gate. Teams can run the same scenario against successive application releases and compare a single resiliency score with the underlying outage evidence.

That approach should make regressions easier to distinguish from harmless alert noise. It also broadens the audience for chaos results: platform teams get Prometheus and infrastructure detail, while application teams get endpoint and recovery-time measurements tied to a release.

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.