September 2026Unreviewed
HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation
Nikita Oblakov, Sabrina Sadiekh, Evgeniy Kokuykin
Abstract
Production LLMs must handle inputs that attempt to override system instructions, bypass safety policies or elicit harmful responses. A common mitigation is a separate guardrail model. Existing reports, however, provide little evidence on Russian prompt injection or Russian surface obfuscation. We present HiveTraceGuard-Pro, a 0.6B generative guardrail LoRA-tuned from Qwen3-0.6B. It is trained on Russian and English and uses one binary scoring rule (safe/unsafe) for the final target turn. Its tra
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{oblakov2026hivetraceguardpro,
title = {{HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation}},
author = {Nikita Oblakov and Sabrina Sadiekh and Evgeniy Kokuykin},
year = {2026},
month = sep,
eprint = {2609.01046},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.01046}
}