live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
releaseAI

RamaLama 0.25 makes Toolbox a first-class local AI workspace

The release can drive host container engines and accelerators from Fedora Toolbox while changing model serving from network-visible to local-only by default.

Fedora Toolbox proxies host AI containers and now defaults model serving to loopback.
AI-generated illustration
By The News Desk· Sep 26, 2026the quick take — two AI hosts go live when you do

RamaLama 0.25 turns Fedora Toolbox from an unsupported nesting case into a practical front end for host-side local AI containers. The September 25 release also changes ramalama serve to bind to loopback by default, narrowing an exposure that previously made a served model reachable from the local network without an explicit opt-in.

Together, the changes sharpen RamaLama's workstation role: developers can keep the CLI and its dependencies inside a disposable Toolbox environment while still using the host's Podman or Docker engine and accelerator configuration. The release also makes the safer choice the default when that local model becomes an API endpoint.

Toolbox controls the host runtime

RamaLama previously disabled container mode when it detected that it was running inside Toolbox. The new Toolbox integration instead uses flatpak-spawn --host to proxy container-engine commands to the host.

The implementation does more than forward podman or docker. Host-side checks for NVIDIA, AMD and other accelerators run through the same proxy, and RamaLama reads Container Device Interface configuration through Toolbox's /run/host mount. That matters because merely starting a container from inside Toolbox would not guarantee that the selected inference image received the host GPU configuration.

The resulting workflow separates the development shell from model execution without adding another nested container engine. Teams can standardize a Toolbox image for the RamaLama CLI and related tools, while keeping images, model-serving containers and device access under the host engine. Administrators should still treat the host proxy as a trust boundary: a process in that Toolbox can ask the host engine to create containers.

Local means loopback unless users opt out

Before 0.25, the served-model port was published on 0.0.0.0 or [::] by default. The new binding behavior defaults to 127.0.0.1; network exposure now requires an explicit --host 0.0.0.0 or --host ::.

RamaLama handles Podman Machine, Docker Desktop, macOS, Windows and WSL2 separately so the host-side proxy can still make the endpoint reachable on host loopback without widening it to the LAN. Generated Compose and Quadlet configurations inherit the selected bind host.

The change can break an existing client on another machine—or another container—that relied on the old wildcard default. That is intentional compatibility friction: operators now have to declare external reachability and can put authentication, firewall rules or a reverse proxy around it rather than discovering that a local model server was already listening on every interface.

RAG helpers no longer require that exception. A related private-network change joins helper and ingestion containers to a dedicated container network and addresses them by name, while publishing readiness endpoints only on host loopback.

Developers upgrading should test remote clients, generated service definitions and GPU selection. The release includes additional WSL2 and device-passthrough fixes, but the key migration question is simpler: if a model endpoint truly needs to leave the workstation, make that exposure explicit.

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.