Skip to content
Search
paperAugust 2026Unreviewed

aiXamine: Unified Black-Box Evaluation of Cross-Dimensional Trade-offs in LLM Safety, Security, and Privacy

Fatih Deniz, Yazan Boshmaf, Dorde Popovic, Issa Khalil

Abstract

The critical failure modes in deployed large language models (LLMs) are cross-dimensional: a model can score 99.3 in safety alignment while refusing one in three benign queries, or improve across every capability metric while losing 21 points in privacy. Existing evaluation frameworks that assess safety, security, and privacy independently cannot detect these patterns. We introduce aiXamine, a unified black-box platform that evaluates LLM trustworthiness across safety, security, and privacy as i

Categories

Cite

@misc{deniz2026aixamine,
  title = {{aiXamine: Unified Black-Box Evaluation of Cross-Dimensional Trade-offs in LLM Safety, Security, and Privacy}},
  author = {Fatih Deniz and Yazan Boshmaf and Dorde Popovic and Issa Khalil},
  year = {2026},
  month = aug,
  eprint = {2608.20554},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.20554}
}