August 2026Unreviewed
LoRAScan: Detecting Backdoor Prompts in Low-Rank Adapters for Large Language Models via Down-Projection Activation Spikes
Doniyorkhon Obidov, H. Yu, Xiaolong Guo, Kaichen Yang
Abstract
Low-rank adaptation (LoRA) enables efficient specialization and distribution of large language models through compact adapters. However, untrusted adapters introduce a supply-chain threat: a backdoored adapter can cause a model to generate harmful content, malicious code, political propaganda, or covert advertisements when an input contains a hidden trigger. Adapter-agnostic defenses merge the adapter with the base model, which dilutes backdoor signals and reduces detection performance. Existing
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{obidov2026lorascan,
title = {{LoRAScan: Detecting Backdoor Prompts in Low-Rank Adapters for Large Language Models via Down-Projection Activation Spikes}},
author = {Doniyorkhon Obidov and H. Yu and Xiaolong Guo and Kaichen Yang},
year = {2026},
month = aug,
eprint = {2608.06795},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/449779b572b8bdb28df60a03ccc7c18a432b738c}
}