August 2026Unreviewed
aiXamine: Unified Black-Box Evaluation of Cross-Dimensional Trade-offs in LLM Safety, Security, and Privacy
Fatih Deniz, Yazan Boshmaf, Dorde Popovic, Issa Khalil
Abstract
The critical failure modes in deployed large language models (LLMs) are cross-dimensional: a model can score 99.3 in safety alignment while refusing one in three benign queries, or improve across every capability metric while losing 21 points in privacy. Existing evaluation frameworks that assess safety, security, and privacy independently cannot detect these patterns. We introduce aiXamine, a unified black-box platform that evaluates LLM trustworthiness across safety, security, and privacy as i
Categories
Cite
@misc{deniz2026aixamine,
title = {{aiXamine: Unified Black-Box Evaluation of Cross-Dimensional Trade-offs in LLM Safety, Security, and Privacy}},
author = {Fatih Deniz and Yazan Boshmaf and Dorde Popovic and Issa Khalil},
year = {2026},
month = aug,
eprint = {2608.20554},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.20554}
}