Skip to content
Search
paperJanuary 2026Unreviewed

Continual Red-Teaming and Guardrail Distillation for Tool-Using LLM Agents: Prompt-Injection Resistance with Utility Preservation

Wesley Gao

Stout in Computer Science and Technology Studies

Abstract

Tool-using language-model agents can convert indirect prompt injection into consequential actions, making guardrail quality a joint security, utility, and efficiency problem. This study evaluates a ReAct-style control, native tool filtering, deterministic self-verification, a calibrated action gate, and a distilled guardrail on AgentDojo v0.1.22. The evaluation covers 97 benign tasks, 629 canonical attack cases, ten attack formulations, and 32,456 recorded actions, of which 20,953 were assigned

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@article{gao2026continual,
  title = {{Continual Red-Teaming and Guardrail Distillation for Tool-Using LLM Agents: Prompt-Injection Resistance with Utility Preservation}},
  author = {Wesley Gao},
  year = {2026},
  month = jan,
  journal = {Stout in Computer Science and Technology Studies},
  doi = {10.61424/hewtat90},
  url = {https://doi.org/10.61424/hewtat90}
}