RamaLama 0.25 makes Toolbox a first-class local AI workspace
The release can drive host container engines and accelerators from Fedora Toolbox while changing model serving from network-visible to local-only by default.
RamaLama 0.25 turns Fedora Toolbox from an unsupported nesting case into a practical front end for host-side local AI containers. The September 25 release also changes ramalama serve to bind to loopback by default, narrowing an exposure that previously made a served model reachable from the local network without an explicit opt-in.
Together, the changes sharpen RamaLama's workstation role: developers can keep the CLI and its dependencies inside a disposable Toolbox environment while still using the host's Podman or Docker engine and accelerator configuration. The release also makes the safer choice the default when that local model becomes an API endpoint.
Toolbox controls the host runtime
RamaLama previously disabled container mode when it detected that it was running inside Toolbox. The new Toolbox integration instead uses flatpak-spawn --host to proxy container-engine commands to the host.
The implementation does more than forward podman or docker. Host-side checks for NVIDIA, AMD and other accelerators run through the same proxy, and RamaLama reads Container Device Interface configuration through Toolbox's /run/host mount. That matters because merely starting a container from inside Toolbox would not guarantee that the selected inference image received the host GPU configuration.
The resulting workflow separates the development shell from model execution without adding another nested container engine. Teams can standardize a Toolbox image for the RamaLama CLI and related tools, while keeping images, model-serving containers and device access under the host engine. Administrators should still treat the host proxy as a trust boundary: a process in that Toolbox can ask the host engine to create containers.
Local means loopback unless users opt out
Before 0.25, the served-model port was published on 0.0.0.0 or [::] by default. The new binding behavior defaults to 127.0.0.1; network exposure now requires an explicit --host 0.0.0.0 or --host ::.
RamaLama handles Podman Machine, Docker Desktop, macOS, Windows and WSL2 separately so the host-side proxy can still make the endpoint reachable on host loopback without widening it to the LAN. Generated Compose and Quadlet configurations inherit the selected bind host.
The change can break an existing client on another machine—or another container—that relied on the old wildcard default. That is intentional compatibility friction: operators now have to declare external reachability and can put authentication, firewall rules or a reverse proxy around it rather than discovering that a local model server was already listening on every interface.
RAG helpers no longer require that exception. A related private-network change joins helper and ingestion containers to a dedicated container network and addresses them by name, while publishing readiness endpoints only on host loopback.
Developers upgrading should test remote clients, generated service definitions and GPU selection. The release includes additional WSL2 and device-passthrough fixes, but the key migration question is simpler: if a model endpoint truly needs to leave the workstation, make that exposure explicit.
sources
- RamaLama v0.25.0 releasegithub.com
- RamaLama Toolbox support pull requestgithub.com
- RamaLama loopback-default pull requestgithub.com
- RamaLama private RAG network pull requestgithub.com
- RamaLama GPU and WSL2 fixes pull requestgithub.com
comments · 0