September 2026Unreviewed
LLM Forensics: Where Do Backdoors Hide? Localizing and Controlling Trigger Mechanisms with Sparse Autoencoders
Wissam Antoun, Francis Kulumba, Théo Lasnier, Benoît Sagot, Djamé Seddah
Abstract
Even though backdoors in LLMs have been a growing concern, their inner workings are still under heavy scrutiny. Trigger-based backdoors are easy to define behaviorally, a rare input that makes the model switch to a chosen response pattern, but the mechanism between triggers and their responses is less clear. We study this mechanism in a controlled, harmless language-switching setting, where fixed trigger sequences make 1B and 8B language models continue English prompts in French or German. For t
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{antoun2026llm,
title = {{LLM Forensics: Where Do Backdoors Hide? Localizing and Controlling Trigger Mechanisms with Sparse Autoencoders}},
author = {Wissam Antoun and Francis Kulumba and Théo Lasnier and Benoît Sagot and Djamé Seddah},
year = {2026},
month = sep,
eprint = {2609.07746},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.07746}
}