Garak gives local vLLM teams a first red-team gate—but not a safety verdict
Red Hat’s walkthrough turns one prompt-injection probe into reproducible evidence while showing why a pass/fail label needs context.
Red Hat has published a reproducible first red-team workflow for an open-weight model: serve IBM Granite 4.1 3B through vLLM, point NVIDIA’s garak scanner at its OpenAI-compatible endpoint and inspect the prompts that defeated the model.
The value is not the tutorial’s pass/fail label. It is the shift from an informal jailbreak conversation to a recorded test that teams can repeat after changing a model, system prompt or serving stack.
The smallest useful scan
The walkthrough installs garak in a Python virtual environment and serves Granite locally with vLLM. It then configures garak’s OpenAI-compatible target and runs the promptinject.HijackHateHumans probe.
Garak sends adversarial prompts, captures responses and applies detectors that score whether the model followed the attack. Red Hat’s example processed 1,280 attempts and reported an 83.36% attack success rate for the selected detector. That number applies to this probe, configuration and model run; it is not a general safety score for Granite.
The resulting JSONL, HTML summary and hit log are more useful than the headline percentage. They let engineers inspect individual failures, preserve evidence and compare later runs.
What the result does—and does not—mean
A garak failure shows that a model complied with prompts matching the selected detector. It does not prove that a deployed application is exploitable in the same way, because production behavior also depends on system prompts, retrieval data, tool permissions, output handling and controls around the model.
Conversely, passing one probe does not make a model safe. Garak includes probes for instruction hijacking, jailbreak personas, encoding tricks, malicious signatures and cross-site-scripting output. Human testers still matter for application-specific attacks that a generic corpus will not anticipate.
Teams should therefore treat the first scan as a baseline. Record the model identifier, serving version, decoding settings, probe set and detector configuration. Then rerun the same suite when any of those inputs change.
Put evidence into the delivery gate
The next step is CI integration, but a single universal threshold would be misleading. A better policy separates regressions from absolute risk: block a build when previously resisted prompts begin succeeding, and route high-impact failures for review even if the aggregate score improves.
The application boundary matters most for agentic systems. A prompt that produces disallowed text is different from one that causes a tool call with infrastructure privileges. Test the model with the same tool schema and authorization boundary it will have in production, and keep execution controls outside the model.
Garak makes adversarial testing accessible on hardware teams already own. Its output becomes valuable when it is versioned, compared and connected to application-level consequences—not when one score is mistaken for certification.
sources
- Red team your AI model with garakdevelopers.redhat.com
comments · 0