Skip to content

Responsible AI

Fairness, bias, transparency, and ethical AI considerations

Resources
13

Newest first · 6 reviewed on this page

Search instead
paper2026Unreviewed

Trustworthy AI LLM Scalability Risk Index (LSRI): A Cybersecurity Framework Assessing Agentic-AI Security & Software Model Supply Chain Safety Boosting AI-Generated Malware Defense & Explainability Mitigating Emerging Risks of Generative AI

Kiarash Ahi, Vaibhav Agrawal, Saeed Valizadeh

As AI shifts from human-in-the-loop interfaces to autonomous multi-agent systems capable of real-time code execution and tool integration through protocols like the Model Context Protocol (MCP), traditional SAST, DAST, and legacy AI safety methods fail to detect modern…

paper2026Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2Unreviewed

Shifting the Unit of Safety: From Model to System in the Generative and Agentic Era

Sakshi Jain

For a decade, responsible AI at internet scale rested on a reassuring assumption: that risk lives primarily in a model, it is a unit you can isolate, and that privacy, fairness and safety is therefore something you certify at model level, before launch. At LinkedIn, where AI…

paper2026Unreviewed

Artificial Intelligence Security and Adversarial Machine Learning: Threat Models, Defensive Strategies, and Forensic Implications for Trustworthy AI Systems

James H. Senanu

The rapid integration of artificial intelligence (AI) systems into security-critical domains has introduced new vulnerabilities, exposing these systems to a growing spectrum of adversarial threats. Adversarial machine learning (AML) has emerged as a key area of research aimed at…

paper2022arXiv preprintReviewed

Constitutional AI: Harmlessness from AI Feedback

Yuntao Bai, Saurav Kadavath, Sandipan Kundu +44

Introduces Constitutional AI (CAI), a method for training AI systems to be harmless using a set of principles (a constitution) and AI-generated feedback, reducing reliance on human red teamers.

GuardrailsResponsible AIOpen access1,100 cit.