June 2026Unreviewed
BadAgent Extension
Pedro Yanes Garrido, Diego Fernandez Arias
Abstract
This paper presents an empirical study on backdoor attacks in large language model agents. We extend a recent attack framework by adding two lightweight benchmarks that measure cross-domain robustness and trigger visibility without changing the model architecture. Our approach fine-tunes AgentLM-based agents with parameter-efficient methods on operating system and web browsing tasks using multiple poisoning ratios and both visible and invisible triggers. We then evaluate the agents with
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{garrido2026badagent,
title = {{BadAgent Extension}},
author = {Pedro Yanes Garrido and Diego Fernandez Arias},
year = {2026},
month = jun,
doi = {10.4018/979-8-3373-8252-4.ch010},
url = {https://doi.org/10.4018/979-8-3373-8252-4.ch010}
}