Skip to content
Search
paperSeptember 2026Unreviewed

HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation

Nikita Oblakov, Sabrina Sadiekh, Evgeniy Kokuykin

Abstract

Production LLMs must handle inputs that attempt to override system instructions, bypass safety policies or elicit harmful responses. A common mitigation is a separate guardrail model. Existing reports, however, provide little evidence on Russian prompt injection or Russian surface obfuscation. We present HiveTraceGuard-Pro, a 0.6B generative guardrail LoRA-tuned from Qwen3-0.6B. It is trained on Russian and English and uses one binary scoring rule (safe/unsafe) for the final target turn. Its tra

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{oblakov2026hivetraceguardpro,
  title = {{HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation}},
  author = {Nikita Oblakov and Sabrina Sadiekh and Evgeniy Kokuykin},
  year = {2026},
  month = sep,
  eprint = {2609.01046},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.01046}
}