September 2026Unreviewed
EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents
Nanxi Li, Yingzi Ma, Yulong Cao, Edward Suh, Bo Li, Dawn Song, Chaowei Xiao
Abstract
Large Language Model (LLM) agents are turning language into real-world effects, making safety necessary against both indirect prompt injections and direct harmful requests. System-level safety harnesses add an enforcement layer beyond model-level defenses, but existing harnesses are usually designed once by experts and applied across heterogeneous models and domains. Effective protection is deployment-dependent: models differ in how much enforcement they need before utility declines, while domai
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
Suggested from the entry's categories.
Cite
@misc{li2026evosafeharness,
title = {{EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents}},
author = {Nanxi Li and Yingzi Ma and Yulong Cao and Edward Suh and Bo Li and Dawn Song and Chaowei Xiao},
year = {2026},
month = sep,
eprint = {2609.05903},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.05903}
}