Skip to content

Tool Use Security

Function calling security, plugin safety, and API tool governance

Resources
44

Newest first · 9 reviewed on this page

Search instead
paper2026Journal of Computer Virology and Hacking TechniquesUnreviewed

Securing LLM-based agents against cyberattacks: a comprehensive survey on attack techniques and defense strategies

Nyashadzashe Tamuka, T. Mathonsi, T. Olwal +3

Large Language Model (LLM)-based agents integrate various models, including planning loops, memory, tool use, and multi-agent systems, enabling autonomous decision-making through natural-language interfaces. This autonomy also expands the cyberattack surface from model-only…

paper2026Unreviewed

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks

P. Wang, Ao-Jie Yuan, Haiyu Zhang +3

Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on the user's behalf, maintains state, and communicates with other agents through social and task…

paper2026Unreviewed

When Agents Overtrust Environmental Evidence: An Extensible Agentic Framework for Benchmarking Evidence-Grounding Defects in LLM Agents

Strick Sheng, Ziyue Wang, Liyi Zhou

Large language model agents increasingly operate through environment-facing scaffolds that expose files, web pages, APIs, and logs. These observations influence tool use, state tracking, and action sequencing, yet their reliability and authority are often uncertain.…

paper2026Unreviewed

Agentic Microphysics: A Manifesto for Generative AI Safety

Federico Pierucci, Matteo Prandi, Marcantonio Bracale Syrnikov +2

This paper advances a methodological proposal for safety research in agentic AI. As systems acquire planning, memory, tool use, persistent identity, and sustained interaction, safety can no longer be analysed primarily at the level of the isolated model. Population-level risks…

paper2026Unreviewed

The Gate Is Only as Honest as Its Contracts: ContractGuard for the Contract Layer of Risk-Aware Causal Gating

Laxmipriya Ganesh Iyer, Rahul Suresh Babu

Risk-Aware Causal Gating (RACG) defends tool-augmented LLM agents against indirect prompt injection by removing dangerous tools from the agent's visible action space, so that even a fully injection-compliant agent cannot call a tool it cannot see. We make three points. First,…

paper2026Unreviewed

Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems

Jimmy Laurence Rippin, Simon C. Marshall, David Demitri Africa +1

Increasingly autonomous agentic AI systems pose novel multi-agent risks, such as secret collusion via covert communication channels. The natural defence to these collusion attempts is to monitor plain-text communication, but the efficacy of monitors has been called into doubt by…

paper2026Unreviewed

NetInjectBench: Benchmarking Indirect Prompt Injection in Tool-Using Large Language Model Agents for Network Operations

Ruksat Khan Shayoni, Muhammad Faraz Shoaib, S M Asif Hossain +1

Tool-using large language model (LLM) agents are attractive for network operations, but tickets, alerts, logs, runbooks, and ChatOps messages can carry indirect prompt injections. We present NetInjectBench, a 130-scenario benchmark that separates untrusted artifact text, trusted…

paper2026Unreviewed

Engineering Trustworthy Agentic AI for Critical Systems

Omar Al-Refai, Ibrahim Shahbaz, Adam Ali Husseinat +3

Agentic artificial intelligence systems, capable of autonomous perception, planning, tool use, and multi-step action, are increasingly proposed for critical engineering domains where decisions carry physical, operational, or economic consequences. This survey addresses a gap in…

paper20262026 International Conference on Connected Intelligence for Industrial Applications (CI2A)Unreviewed

AttaX-Multimodal: End-to-End Evaluation of Multimodal AI Safety with Threatscore and Resiliencescore

Rahul Karne

Currently, there is no single benchmark that can be used to measure the safety and resilience of multimodal AI assistants when subjected to malicious attacks. To fill this gap, we have created AttaX-Multimodal, a comprehensive benchmark of multimodal AI assistant safety that…

paper2026Unreviewed

Influence Is Not Authority: When Causal Guardrail Signals Make Legitimate Tool Use Look Like an Attack in Tool-Using LLM Agents

Tanzim Ahad, Ismail Hossain, Md Jahangir Alam +3

The key limitation of current state-of-the-art influence-based guardrails is that they do not reliably distinguish a legitimate, user-authorized action from a malicious, unauthorized action when both rely on external tool information. This ambiguity can cause benign actions to…

paper2026Unreviewed

Skynet: Workflow-Level Anomaly Detection for Agentic AI via Semantic and Structural Modeling

Chaoyu Zhang, Hexuan Yu, Heng Jin +6

Agentic AI systems execute complex tasks through long-horizon workflows of planning, tool use, and multi-agent coordination. Task failures in these systems often originate from a single step, such as an injected prompt or a flawed plan, and are then amplified through downstream…

paper2026Unreviewed

When Malicious Instructions Persist: Persistent Memory Poisoning Attack on Harness-Based Agents

Shuhuai Huang, Jingfeng Zhang, Hong Jia

Harness design has transformed the development of LLM-based agents by integrating memory, tool use, and runtime control. However, this design also introduces security and privacy risks because malicious instructions from external sources may be written into persistent memory and…

paper2026CEUR Workshop Proceedings, Vol-4260: Proceedings of the 8th Workshop for Young Scientists in Computer Science & Software Engineering (CS&SE@SW 2025)Unreviewed

A framework for efficient and secure LLM agency: a case for the GraphQL paradigm

Viktor Zhakhalov

LLM agents must translate natural language into concrete actions on external tools. Most systems use JSON-based function calling or, more riskily, let models emit imperative code. We propose a GraphQL-first alternative that reframes tool use as typed, declarative program…

paper2026International Journal For Multidisciplinary ResearchUnreviewed

Adversarial Robustness of Foundation Models for Intelligent Mechanical Systems: Threat Models, Benchmarks, and Defense Stacks

Vishwanath

Foundation models increasingly operate across modalities (vision, language, audio, and vision–language) and are deployed in decision-critical pipelines with tool use and retrieval. This expands the adversarial surface: small perturbations to images or audio can flip predictions,…

paper2026Unreviewed

Prompt Injection Attacks Against Clinical LLM Agents Accessing Electronic Health Records: A Survey, Threat Model, Benchmark Specification, and Layered Defense Synthesis

Divya Pandey, Shivani Manchanda, Gangesh Pathak +1

Clinical large language model (LLM) agents are entering production hospital deployments, where they read longitudinal electronic health records (EHRs), retrieve evidence from clinical knowledge bases, and assist with summarization, dosing, triage, and guideline-based decisions.…