May 2026Unreviewed
When Efficiency Backfires: Cascading LLMs Trigger Cascade Failure under Adversarial Attack
Zehan Sun, Dingfan Chen, Songze Li
Abstract
Large Language Model (LLM) cascade systems are designed to balance efficiency and performance by processing queries with lightweight models while selectively escalating complex cases to more powerful ones. Such systems seek to reduces computational cost and latency while maintaining task performance, making it an appealing choice for large-scale deployment. However, the cascade design introduces new vulnerabilities through an expanded attack surface: the inclusion of lightweight front-end models
Categories
Framework mappings
MITRE ATLAS
- AML.T0043Craft Adversarial Data
Suggested from the entry's categories.
Cite
@misc{sun2026when,
title = {{When Efficiency Backfires: Cascading LLMs Trigger Cascade Failure under Adversarial Attack}},
author = {Zehan Sun and Dingfan Chen and Songze Li},
year = {2026},
month = may,
eprint = {2605.17288},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.17288}
}