October 2024UnreviewedOpen access
Mitigating adversarial manipulation in LLMs: a prompt-based approach to counter Jailbreak attacks (Prompt-G)
Bhagyajit Pingua, Deepak Murmu, Meenakshi Kandpal, Jyotirmayee Rautaray, Pranati Mishra, R. Barik, Manob Saikia
PeerJ Computer Science
Abstract
Large language models (LLMs) have become transformative tools in areas like text generation, natural language processing, and conversational AI. However, their widespread use introduces security risks, such as jailbreak attacks, which exploit LLM’s vulnerabilities to manipulate outputs or extract sensitive information. Malicious actors can use LLMs to spread misinformation, manipulate public opinion, and promote harmful ideologies, raising ethical concerns. Balancing safety and accuracy require
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
- LLM02Sensitive Information Disclosure
MITRE ATLAS
- AML.T0024.000Infer Training Data Membership
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@article{pingua2024mitigating,
title = {{Mitigating adversarial manipulation in LLMs: a prompt-based approach to counter Jailbreak attacks (Prompt-G)}},
author = {Bhagyajit Pingua and Deepak Murmu and Meenakshi Kandpal and Jyotirmayee Rautaray and Pranati Mishra and R. Barik and Manob Saikia},
year = {2024},
month = oct,
journal = {PeerJ Computer Science},
doi = {10.7717/peerj-cs.2374},
url = {https://www.semanticscholar.org/paper/a5835f37febb7b65e56e27c07beb1b9fa5a1bea4}
}