August 2026Unreviewed
An Empirical Study of Output-to-Input Loops for Black-Box Backdoor Detection in Fine-Tuned Open-Weight LLMs
Md Nahid Hasan, Mohammad Arif Hossain
Abstract
Anyone can upload a fine-tuned large language model (LLM) to a public repository and claim it is safe. A backdoored model behaves normally on ordinary inputs until a hidden trigger fires, and a user with no training data, clean reference weights, or the trigger phrase has no clear way to check the model before using it. We introduce and empirically evaluate self-feeding, a black-box test method that feeds a model's own output back as its next input, so the text drifts away from the starting prom
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{hasan2026empirical,
title = {{An Empirical Study of Output-to-Input Loops for Black-Box Backdoor Detection in Fine-Tuned Open-Weight LLMs}},
author = {Md Nahid Hasan and Mohammad Arif Hossain},
year = {2026},
month = aug,
eprint = {2608.11348},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/332aa7504f08004ae724a19f94043c1278e21ce4}
}