Skip to content
Search
paperMay 2026Unreviewed

When Efficiency Backfires: Cascading LLMs Trigger Cascade Failure under Adversarial Attack

Zehan Sun, Dingfan Chen, Songze Li

Abstract

Large Language Model (LLM) cascade systems are designed to balance efficiency and performance by processing queries with lightweight models while selectively escalating complex cases to more powerful ones. Such systems seek to reduces computational cost and latency while maintaining task performance, making it an appealing choice for large-scale deployment. However, the cascade design introduces new vulnerabilities through an expanded attack surface: the inclusion of lightweight front-end models

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data

Suggested from the entry's categories.

Cite

@misc{sun2026when,
  title = {{When Efficiency Backfires: Cascading LLMs Trigger Cascade Failure under Adversarial Attack}},
  author = {Zehan Sun and Dingfan Chen and Songze Li},
  year = {2026},
  month = may,
  eprint = {2605.17288},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.17288}
}