Skip to content
Search
paperAugust 2026Unreviewed

LoRAScan: Detecting Backdoor Prompts in Low-Rank Adapters for Large Language Models via Down-Projection Activation Spikes

Doniyorkhon Obidov, H. Yu, Xiaolong Guo, Kaichen Yang

Abstract

Low-rank adaptation (LoRA) enables efficient specialization and distribution of large language models through compact adapters. However, untrusted adapters introduce a supply-chain threat: a backdoored adapter can cause a model to generate harmful content, malicious code, political propaganda, or covert advertisements when an input contains a hidden trigger. Adapter-agnostic defenses merge the adapter with the base model, which dilutes backdoor signals and reduces detection performance. Existing

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{obidov2026lorascan,
  title = {{LoRAScan: Detecting Backdoor Prompts in Low-Rank Adapters for Large Language Models via Down-Projection Activation Spikes}},
  author = {Doniyorkhon Obidov and H. Yu and Xiaolong Guo and Kaichen Yang},
  year = {2026},
  month = aug,
  eprint = {2608.06795},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/449779b572b8bdb28df60a03ccc7c18a432b738c}
}