Skip to content
Search
paperAugust 2026Unreviewed

An Empirical Study of Output-to-Input Loops for Black-Box Backdoor Detection in Fine-Tuned Open-Weight LLMs

Md Nahid Hasan, Mohammad Arif Hossain

Abstract

Anyone can upload a fine-tuned large language model (LLM) to a public repository and claim it is safe. A backdoored model behaves normally on ordinary inputs until a hidden trigger fires, and a user with no training data, clean reference weights, or the trigger phrase has no clear way to check the model before using it. We introduce and empirically evaluate self-feeding, a black-box test method that feeds a model's own output back as its next input, so the text drifts away from the starting prom

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{hasan2026empirical,
  title = {{An Empirical Study of Output-to-Input Loops for Black-Box Backdoor Detection in Fine-Tuned Open-Weight LLMs}},
  author = {Md Nahid Hasan and Mohammad Arif Hossain},
  year = {2026},
  month = aug,
  eprint = {2608.11348},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/332aa7504f08004ae724a19f94043c1278e21ce4}
}