paper/2026Unreviewed
Viet K. Nguyen, Mohammad I. Husain
Agentic AI frameworks let a language model plan, keep memory, and call tools that reach real files, mail, and services. Most of these agents also read images, which gives an attacker a way to put text into the agent's context without going through the user. We present MMPIBench,…
paper/2026Unreviewed
Anna Gazani, Spyridon Kounoupidis, Panagiotis Katsaros +3
The integration of Large Language Models (LLMs) into Security Operations Centers (SOCs) streamlines threat intelligence but introduces critical vulnerabilities, notably indirect prompt injection via log poisoning. Adversaries exploit this vector to execute multistep…
paper/2026Unreviewed
Zehua Zhang, Jie Hu, Pratham Hegde +13
Conventional vulnerability analysis relies on either system access or dynamic interaction, all of which may be unavailable to third-party analysts auditing closed-source, remotely hosted, critical in situ systems, or commercially gated software. Therefore, we propose a new…
paper/2026Unreviewed
Asif Pinjari, Mithun Paul Saint-Germain
When an indirect prompt injection succeeds against an LLM agent, the compromise is visible in the agent's own behavior: a benign prefix of tool calls, a poisoned observation, and a suffix of actions that serve the attacker. An operator needs three facts: where the attack…
paper/2026Unreviewed
Lisa Bouger, Yannick Teglia, Philippe Loubet Moundi
We propose an influence score to quantify the contribution of attention heads to classification decisions in Transformer-based models designed for prompt injection detection. The score combines directional influence on the logits with structural contribution within the residual…
paper/2026International Journal of Information SecurityUnreviewed
Sebastián Vargas Yáñez, Sergio Tobón
Public Large Language Model (LLM) APIs draw attacker traffic that defenders cannot see. Probes hit at the semantic layer—past TLS, past Web Application Firewall rules—and conventional intrusion detection picks up almost none of it. No open-source honeypot framework today…
paper/2026Yalvaç akademi dergisiUnreviewed
Oğuzhan Kilim
The widespread adoption of systems based on Large Language Models has made the reliable detection of prompt injection attacks a critical requirement. However, high performance achieved on training and test splits generated from the same data source does not guarantee that models…
paper/2026Unreviewed
Suyoung Lee, Myungsub Choi
Verdict-only evaluation does not reveal whether a vision-language model (VLM) used the visual evidence that should support its decision. We study this problem in web-agent guardrails, where a VLM judges whether on-screen text conflicts with a user instruction. We introduce…
paper/2026IEEE Transactions on Dependable and Secure ComputingUnreviewed
Yuchen Yang, Yi-Ming Li, H. Yao +5
Recent advancements have led to the widespread adoption of code-oriented large language models (Code LLMs) for programming tasks. Despite their success in deployment, their security research is left far behind. This paper introduces a new attack paradigm: (automatic) external…
paper/2026Security and PrivacyUnreviewed
Albert Adusei Brobbey, Narayan P. Bhosale
Agentic artificial intelligence systems increasingly depend on persistent runtime memory, including vector databases, episodic memory stores, long‐term retrieval indices, and cloud‐to‐edge replicas. Existing security frameworks address prompt injection, data poisoning, and…
paper/2026IEEE Transactions on Dependable and Secure ComputingUnreviewed
Wen-Jing Chen, Jie Cui, Wenjie Huang +3
While Large Language Models (LLMs) have achieved revolutionary advancements in natural language processing, their inherent vulnerability to prompt injection attacks has raised significant security concerns. Existing security detection approaches for LLM deployment often fail to…
paper/2026ComputersUnreviewed
Adil Khan, Khaled AlKhanbashi, Azza Mohamed
Large language model (LLM) agents that retrieve external content and use tools are vulnerable to indirect prompt injection, in which untrusted content contains instructions intended to influence agent behavior. We evaluated four defenses and an undefended control across GPT-5.4,…
paper/2026Unreviewed
Wu-Jie Xiong, Rabimba Karanjai, Yang Lu +2
Large language model agents place outputs from external skills into their execution context, allowing attacker-controlled data to influence later privileged actions. Existing defenses mainly classify untrusted content or authorize proposed operations. They do not directly…
paper/2026Unreviewed
Ziliang Zhang, Yubo Zhu, Wei Tong +4
Retrieval-Augmented Generation (RAG) has emerged as a powerful paradigm for combining large language models (LLMs) with external knowledge sources. However, RAG systems remain vulnerable to prompt injection attacks, which may mislead the retriever or generator to expose…
paper/2026ElectronicsUnreviewed
Zhaowen Feng, Zhenhui Liu, Ming-Jun Ma +2
Tool-level attacks on Large Language Model (LLM) agents—poisoned tool descriptions, prompt injection, and capability misrepresentation—are universally effective, yet no existing defense provides comprehensive protection. We propose Architectural Intent Collapse (AIC), a formal…
paper/2026ElectronicsUnreviewed
Sana Mourad, E. E. Abdallah, Mohammad Ababneh
Current language model deployments face growing security challenges from prompt-based attacks, including jailbreaks, direct and indirect prompt injection, and instruction hijacking, which often evade traditional rule-based safeguards. As these models are increasingly integrated…
paper/2026FinTechUnreviewed
Chong-Hui Tan, Qinxu Ding
The rapid adoption of large language models (LLMs) in financial services has generated a growing literature on “responsible AI” in domains such as investment analysis, credit assessment, risk management, compliance, and financial advisory systems. Unlike earlier AI systems, LLMs…
paper/2026Unreviewed
Jiawei Liu, Jiacheng Guo, Tian Zhang +4
Foundation models are increasingly used for perception, reasoning, planning, and action generation in embodied agents, creating security risks that can propagate from digital inputs to physical behavior. Existing surveys often organize threats by mechanisms such as jailbreaks,…
paper/2026ElectronicsUnreviewed
Alaa Alnemari, Mashael M. Alsulami
Large Language Models (LLMs) are increasingly deployed in security-critical applications but remain vulnerable to indirect prompt injection attacks that cannot be fully addressed by conventional prompt detection techniques. This paper proposes DT-GenShield, a Digital Twin-driven…
paper/2026Applied SciencesUnreviewed
Wonbae Kim, Hee-Kyong Yoo, Nammee Moon
The deployment of Large Language Model (LLM)-generated SQL in Artificial Intelligence of Things (AIoT) systems introduces critical security risks, as prompt injection attacks can manipulate LLMs into producing unauthorized queries that expose sensitive data or execute…
paper/2026ElectronicsUnreviewed
Adam Ait Hsine, A. Arabo
The deployment of large language models (LLMs) in real-world applications introduces a compounding security problem: detecting adversarial inputs such as prompt injection and jailbreak-driven data leakage while simultaneously preventing the detection mechanism itself from…
paper/2026Al-Noor Journal of Engineering Management and Computer ScienceUnreviewed
Fatimah Alhamzawi
Large language models (LLMs) embedded in enterprise workflows cannot structurally distinguish legitimate instructions from adversarial ones in the same token stream, making prompt injection OWASP's top LLM risk for two consecutive editions a persistent threat across direct and…
paper/2026Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2Unreviewed
Lu Lin, Jinghui Chen, Ting Wang +4
Large language models (LLMs) are increasingly embedded as core components of data-centric systems, supporting analytical decision making, and automated reasoning over large-scale, heterogeneous datasets. Yet their deployment in open-world environments raises fundamental…
paper/2026Applied SciencesUnreviewed
Rodrigo Tavares de Pina Simões, Xavier Larriva-Novo, Carmen Sánchez-Zas +2
Security Operations Centers (SOCs) rely on Level 1 analysts to triage increasing alert volumes amid alert fatigue and tool fragmentation. LLM-based multi-agent systems using the Model Context Protocol (MCP) are being adopted to automate these tasks, but their autonomy and tool…
paper/2026Unreviewed
He Zhang, Fei-Long Li, Ding-Ning Long +5
Smart-home assistants increasingly use multimodal large language models (MLLMs) that perceive video and audio directly. This raises a safety question specific to the home: can the agent tell a genuine user command from ambient or externally-sourced content, television speech,…
paper/2026Unreviewed
Jia-Chen Zhang, Zenghui Zhang, Kai-Wei Zhang
Computer-use agents (CUAs), which empower large language models to autonomously operate operating systems and the web, are increasingly vulnerable to indirect prompt injection attacks. A widely adopted defense is the human-in-the-loop paradigm, in which the agent pauses for…
paper/2026Scientific Journal of Computer ScienceUnreviewed
Victor Omoboye Oluwasegun, O. Falebita, N. Adebola +6
The growing cybersecurity vulnerabilities in artificial intelligence (AI) service models, particularly Large Language Models (LLMs), highlight code injection as a critical threat to chatbot reliability and safe deployment. On the account that LLMs process inputs as…
paper/20262026 International Conference on Intelligent Multimedia, Networking, and Security (IMNS)Unreviewed
Shazid Bin Zaman, Sohan Gyawali, C. Popoviciu +2
Smart home virtual assistants are increasingly powered by large language models to enable information retrieval and home device actuation. As a result, intelligent home environments are becoming more exposed to untrusted inputs, increasing their susceptibility to prompt…
paper/2026British Journal of AnaesthesiaUnreviewed
A. de Cassai, B. Dost, E. Pistollato +3
paper/2026The Scholar Journal for Sciences & TechnologyUnreviewed
Galal Eltayeb, Abdalilah Alhalangy
Abstract But now, given the AI revolution and increased interest in bringing virtual agents and assistants to life banks too are testing LLM-powered AI agents that may assist customers, explain and customize products as well as simplify operational work done by bank employees in…
paper/2026Unreviewed
Gustavo Viana
Large Language Models (LLMs) are increasingly deployed in production systems, yet their prompt-based interaction paradigm introduces a novel attack surface encompassing prompt injection, instruction hijacking, and sensitive data leakage. This paper proposes and empirically…
paper/2026Unreviewed
Nazar Waheed
Abstract Large language model (LLM) applications increasingly combine probabilistic language interpretation with retrieval, persistent memory, external tools, and delegated enterprise credentials. This creates a system-security problem in which untrusted semantic content can…
paper/2026Unreviewed
Swethaa R
Prompt injection attacks pose a critical security threat to large language model (LLM) pipelines, enabling adversaries to hijack model behavior by embedding malicious instructions within user inputs or retrieved tool outputs. Existing defences are primarily rule-based filters…
paper/2026Unreviewed
Tayyeb Nadeem Somro, Bilal Arshad, Ammara Gul
Security Operations Centres (SOCs) are increasingly deploying Large Language Model (LLM) assistants to accelerate threat triage, alert prioritisation, and incident response. While these systems offer substantial productivity gains, their integration into security-critical…
paper/2026Unreviewed
Saurabh Sharma
Abstract Prompt injection—the manipulation of an LLM-based agent through adversarial content embedded in user input—poses a critical security risk in production agentic systems with access to sensitive datastores and code execution environments. Detection-based defenses (keyword…
paper/2026Unreviewed
Jigesh Sheoran
Large Language Model (LLM) APIs are increasingly embedded into production web applications, creating a new class of security vulnerability: prompt injection. An adversarial user can embed instructions within their input that override the application's system prompt, exfiltrate…
paper/2026Unreviewed
Shreya Singh
GraphShield is a graph-structured defense framework for LLMs that represents system prompts, retrieved knowledge, agents, and parsed instructions as directed Trust-Knowledge Graph (TKG). Security is formalized as reachability from an instruction node to policy node in a…
paper/2026International Journal of Advanced Artificial Intelligence ResearchUnreviewed
Grigorii Danileiko
The proliferation of large language model (LLM) systems in legal technology platforms has created a new class of web-interface security vulnerabilities that existing application security frameworks address incompletely. This paper examines prompt injection and insecure output…
paper/2026Unreviewed
Aparnaa Mahalaxmi Arulljothi, Theepan Kumar Gandhi
With LLM agents increasingly deployed to autonomously processed external content like web pages, emails, documents, and API responses become targets of indirect prompt injection attacks where malicious instructions embedded in external content hijack the agent's behavior.…
paper/2026Unreviewed
Min Htet Myet
Abstract Applications built on large language models (LLMs) typically forward user input to a model provider with no enforcement layer in between, leaving prompt-injection attempts and personally identifiable information (PII) to pass through unfiltered in both directions. We…
paper/2026Unreviewed
Yujin kim, Jaekwang Kim
Consumer personas are essential for user understanding and marketing strategy, yet manual construction remains costly and difficult to scale. We propose a data-driven persona generation framework that systematically compares three interpretable text mining methods-SNA, LDA, and…
paper/2026Unreviewed
Vinícius Negrão, Maíra Bocci, Paulo Pitrez
We present FEW-AI-SERIAL, a domain-agnostic semantic compression protocol that reduces structured data payloads by 68-92% while maintaining full human readability and native interpretability by Large Language Models (LLMs). Unlike binary serialization formats (Protocol Buffers,…
paper/2026Unreviewed
Haochuan Wang, Zechen Zhang
Multi-agent LLM systems are entering productionprocessing documents, managing workflows, acting on behalf of users-yet their resilience to prompt injection is still evaluated with a single binary: did the attack succeed? This leaves architects without the diagnostic information…
paper/2026Unreviewed
Divya Pandey, Shivani Manchanda, Gangesh Pathak +1
Clinical large language model (LLM) agents are entering production hospital deployments, where they read longitudinal electronic health records (EHRs), retrieve evidence from clinical knowledge bases, and assist with summarization, dosing, triage, and guideline-based decisions.…
paper/2026Stout in Computer Science and Technology StudiesUnreviewed
Wesley Gao
Tool-using language-model agents can convert indirect prompt injection into consequential actions, making guardrail quality a joint security, utility, and efficiency problem. This study evaluates a ReAct-style control, native tool filtering, deterministic self-verification, a…
paper/2026MathematicsUnreviewed
Zanhong Zheng, Jieming Liang, Mengqin Hu +3
Prompt injection detection is commonly studied as a static offline classification problem, yet deployed LLM systems face evolving attacks and distribution shift after deployment. Static detectors are therefore poorly matched to the threat model, while routing every input to a…
paper/2026Unreviewed
Sabin Adhikari, Roshan Paudel, Dipesh Gautam +4
Prompt injection is a serious threat to the security of large language models operating in AI-powered browsers and autonomous web agents, which depend on the ability of those models to interpret instructions correctly as they are used for automated browsing, data extraction or…
paper/2026Unreviewed
Babak Saravi, Daman Deep Singh, Lara Schorn +4
Abstract Image-embedded prompt injection — adversarial text rendered into the pixel data of medical images — is an emerging threat to vision-language models (VLMs) used in clinical decision support. We systematically evaluated this vulnerability across four production-tier VLMs…
paper/2025Conference on Empirical Methods in Natural Language ProcessingUnreviewed
Hengyu An, Jinghuai Zhang, Tianyu Du +4
Large language model (LLM) agents are widely deployed in real-world applications, where they leverage tools to retrieve and manipulate external data for complex tasks. However, when interacting with untrusted data sources (e.g., fetching information from public websites), tool…
paper/2025arXiv.orgUnreviewed
Chetan Pathade
Large Language Models (LLMs) are increasingly integrated into consumer and enterprise applications. Despite their capabilities, they remain susceptible to adversarial attacks such as prompt injection and jailbreaks that override alignment safeguards. This paper provides a…