Red Hat maps CI tests for agent-skill routing with skill-creator and promptfoo
A Red Hat Developer workflow turns ephemeral trigger checks into durable positive, near-miss and negative tests that can graduate into a blocking CI gate.
Red Hat Developer has laid out a practical testing path for a failure mode that ordinary code review can miss: an agent skill that is valid on disk but does not trigger when users need it — or triggers on the wrong request.
What changed
The engineering guide connects two stages of the skill lifecycle. Anthropic's skill-creator can help authors turn a working workflow into a SKILL.md, generate trigger prompts, grade results and optimize a skill description against a held-out test set. Red Hat's key recommendation is to preserve that work as a repeatable promptfoo suite rather than leave the evaluation inside one authoring session.
The proposed suite has three kinds of cases. Positive prompts should invoke the target skill. Near-miss prompts should route to a sibling skill instead. Negative prompts should invoke neither. Promptfoo's agent SDK providers run those prompts through the actual agent runtime, while its skill-used and not-skill-used assertions verify routing rather than merely checking whether the response mentions a skill by name.
The guide also narrows the test environment deliberately. A fixture working directory lets the agent discover the skill as it would in a real project, instead of reading the source SKILL.md directly and producing a false positive. Discovery can be filtered to the skills under test, and allowed tools can be restricted to read-only operations such as Read, Grep and Glob. That matters because an evaluation run could otherwise push a branch, open a pull request or call an external API.
Who it affects
Teams distributing agent skills across repositories or internal marketplaces have the clearest need for this pattern. Trigger descriptions behave like routing code, but a regression is silent: the general model answers directly and the specialized workflow never runs. Near-miss cases are especially useful when a project has overlapping security, style, release or pre-commit skills.
What to do
Start with the positive and negative prompts produced during authoring, then add near misses for every neighboring skill. Store the promptfoo configuration beside the skill, run it against a minimal fixture, and assert actual skill use. In CI, restrict the job to skill and evaluation-file changes.
Red Hat recommends a progressive rollout: report results without blocking merges first, then turn the suite into a required gate after the team trusts its signal. The article estimates a small trigger suite at roughly $0.10 to $0.15 per evaluation at current prices, though model pricing can change. The larger point is durable regression coverage: treat skill routing as testable behavior, not prose that reviewers can validate by inspection alone.
sources
- Master your skills: Building skills you can trustdevelopers.redhat.com
comments · 0