June 2025Unreviewed
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
Yucheng Li, Surin Ahn, Huiqiang Jiang, Amir H. Abdi, Yuqing Yang, Lili Qiu
arXiv.org
Abstract
Large language models (LLMs) have achieved widespread adoption across numerous applications. However, many LLMs are vulnerable to malicious attacks even after safety alignment. These attacks typically bypass LLMs' safety guardrails by wrapping the original malicious instructions inside adversarial jailbreaks prompts. Previous research has proposed methods such as adversarial training and prompt rephrasing to mitigate these safety vulnerabilities, but these methods often reduce the utility of LLM
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{li2025securitylingua,
title = {{SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression}},
author = {Yucheng Li and Surin Ahn and Huiqiang Jiang and Amir H. Abdi and Yuqing Yang and Lili Qiu},
year = {2025},
month = jun,
eprint = {2506.12707},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2506.12707},
url = {https://www.semanticscholar.org/paper/886194f65d0c2a0b56458d8ba2cd96bd69a97f6a}
}