Skip to content

injection_guard: a labelled set from InjecAgent and AgentDojo #3

Description

@Yunaik

evals/labelled/injection.jsonl holds 20 items and the injection_guard rail scores 20 of 20 on it. A set that size says little about precision and recall on real attacks.

Build a larger labelled set from the public benchmarks in the same JSONL shape (state.tool, state.text, label, note): InjecAgent (1,054 tool-output injections) and AgentDojo (97 tasks, 629 security cases). Then run the rail on it and add the precision and recall table to docs/benchmarks.md:

uv run s1a run injection_guard --slot jev --labelled-set evals/labelled/injection-public.jsonl

docs/roadmap.md § "Prompt-injection guard" has the plan and the two sources. Needs a Jev key (TYPESAFE_API_KEY or OPENROUTER_API_KEY).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions