April 2026Unreviewed
How Code Representation Shapes False-Positive Dynamics in Cross-Language LLM Vulnerability Detection
Maofei Chen, Laifu Wang, Yue Qin, Yuan Wang, Bo Wu, Dongxin Liu
Abstract
How code representation format shapes false positive behaviour in cross-language LLM vulnerability detection remains poorly understood. We systematically vary training intensity and code representation format, comparing raw source text with pruned Abstract Syntax Trees at both training time and inference time, across two 8B-parameter LLMs (Qwen3-8B and Llama 3.1-8B-Instruct) fine-tuned on C/C++ data from the NIST Juliet Test Suite (v1.3) and evaluated on Java (OWASP Benchmark v1.2) and Python (B
Categories
Cite
@misc{chen2026howa,
title = {{How Code Representation Shapes False-Positive Dynamics in Cross-Language LLM Vulnerability Detection}},
author = {Maofei Chen and Laifu Wang and Yue Qin and Yuan Wang and Bo Wu and Dongxin Liu},
year = {2026},
month = apr,
eprint = {2604.27714},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2604.27714}
}