Three prompt-injection defenses move the security decision to different layers
Instruction hierarchy, structural data separation and privilege separation divide responsibility among model vendors, application builders and platform teams.
Red Hat’s Emerging Technologies group frames prompt injection as an architectural problem, not merely a filtering problem. Its Aug. 26 review compares three defenses—an instruction hierarchy, explicit separation of instructions from data, and damage containment—and each puts the hard security decision in a different part of the stack.
That distinction matters for agent builders deciding what they can deploy now. None of the approaches makes prompt injection disappear, and only one can be adopted primarily as an application-architecture choice without waiting for a differently trained model.
Instruction hierarchy depends on the model
Instructional Segment Embedding teaches a model categories such as system instructions, user instructions, data and output. A related instruction-hierarchy technique trains the model to prioritize higher-level instructions and recognize attempts to subvert them.
Red Hat reports improvements over baseline for both approaches, but the results remain incomplete: the best instruction-hierarchy result cited in the review is around 95% successful, with some tests as low as 72%. More importantly for deployment teams, the hierarchy and prioritization logic are established during training or pre-training. An application architect cannot freely redefine those trust rules for every agent workflow.
The practical consequence is dependency. Teams can select models that implement stronger instruction handling and test them against their workloads, but they cannot treat that behavior as an application-controlled authorization boundary.
Instruction/data separation gives applications more control
Architectural Separation of Instructions and Data, or ASIDE, borrows the logic of prepared statements: the application identifies which input is instruction and which is data, while the model represents them differently through fixed rotations of data-token embeddings.
That moves classification closer to the application, where knowledge of the workflow and its trust boundaries already lives. Red Hat describes ASIDE as more robust than asking the model to infer the categories itself. It also says the technique is largely model-agnostic and adds no parameters.
There is still a deployment cost. ASIDE requires an extra fine-tuning step, narrowing the immediately available model choices, and Red Hat says the approach is not perfect. Until model vendors ship compatible conventions, teams cannot bolt this separation onto an arbitrary hosted model through prompt formatting alone.
Containment is deployable, but bespoke
The third family assumes prevention will fail and limits what a compromised agent can do. Red Hat highlights a “three-color” rule: avoid giving one agent all three of untrusted input, sensitive-system access and the ability to change state or exfiltrate data.
Type-directed privilege separation makes that constraint concrete. Unprivileged agents consume risky free-form text and convert it into restricted, pre-approved structured types. Privileged agents act only on that constrained representation.
This is the most directly actionable pattern for today’s agent systems because it is an architecture choice rather than a promise from the model. It is also the least portable. Red Hat notes that teams must design it into each application; removing free-form context can reduce task success, and unprivileged agents can still be driven into denial-of-service behavior that wastes time and tokens.
What to deploy now
The three approaches are complementary, not alternatives with a single winner. Instruction hierarchy can improve model behavior. Structural instruction/data separation can give applications a stronger boundary when compatible fine-tuning is available. Privilege separation can reduce blast radius now, at the price of workflow-specific engineering.
For production agent builders, the review’s clearest implication is to reserve real authority for narrow, typed interfaces and keep untrusted content away from components that can both reach sensitive systems and change state. Guardrails remain useful as detection, but Red Hat’s analysis treats them as a safety net—not the security architecture itself.
sources
- Beyond guardrails: mitigating prompt injection attacksnext.redhat.com
comments · 0