Skip to content
Search
paperNovember 2025Unreviewed

AlignTree: Efficient Defense Against LLM Jailbreak Attacks

Gil Goren, Shahar Katz, Lior Wolf

AAAI Conference on Artificial Intelligence

Abstract

Large Language Models (LLMs) are vulnerable to adversarial attacks that bypass safety guidelines and generate harmful content. Mitigating these vulnerabilities requires defense mechanisms that are both robust and computationally efficient. However, existing approaches either incur high computational costs or rely on lightweight defenses that can be easily circumvented, rendering them impractical for real-world LLM-based systems. In this work, we introduce the AlignTree defense, which enhances mo

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@inproceedings{goren2025aligntree,
  title = {{AlignTree: Efficient Defense Against LLM Jailbreak Attacks}},
  author = {Gil Goren and Shahar Katz and Lior Wolf},
  year = {2025},
  month = nov,
  booktitle = {AAAI Conference on Artificial Intelligence},
  eprint = {2511.12217},
  archivePrefix = {arXiv},
  doi = {10.48550/arXiv.2511.12217},
  url = {https://www.semanticscholar.org/paper/3a45a8e64f082728183c845858a50e5044c73a55}
}