September 2026Unreviewed
Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning
Robin Haselhorst, Lucie Flek, Florian Mai
Abstract
Large language models can hide hidden behaviors that activate only under narrow conditions, such as backdoor triggers, sleeper-agent deployment cues, sandbagging, or topic-conditioned censorship. Such behaviors are difficult to detect without prior knowledge what to look for. We present activation-matched finetuning, an unsupervised detection method that assumes no knowledge of the trigger or the target behavior. Given a suspect model and a publicly available anchor, we finetune the anchor to re
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{haselhorst2026detecting,
title = {{Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning}},
author = {Robin Haselhorst and Lucie Flek and Florian Mai},
year = {2026},
month = sep,
eprint = {2609.00351},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.00351}
}