A new free security checklist provides 222 tests across 20 attack categories to help organizations assess the security of agentic AI systems.
The checklist goes beyond traditional prompt-injection testing. It covers the wider environment around an AI agent, including cloud infrastructure, identity access, connected tools, memory systems, agent-to-agent communication, and deployment pipelines.
The goal is to help security teams identify weaknesses that may exist outside the AI model itself. For example, an exposed MLflow service, accessible cloud credentials, or poor customer-data isolation could create serious security risks even without successfully attacking the underlying AI model.
Security researcher Ravi Rajput developed the checklist using ideas from the OWASP Web Security Testing Guide. It is an independent resource rather than an official OWASP checklist.
20 Areas Covered by the Checklist
The testing process is divided into several stages, starting with understanding what the AI system can access before testing how it responds to malicious instructions.
The checklist covers areas such as:
- Attack surface discovery
- AI orchestration
- Cloud identity and permissions
- Model and software supply chains
- Prompt injection
- System prompt exposure
- Unsafe output handling
- Tool and plugin security
- Excessive agent permissions
- Memory and retrieval systems
- Agent-to-agent communication
- Model Context Protocol (MCP) servers
- Privilege escalation
- Lateral movement
- Persistence
- Data theft
- Resource exhaustion
- Integrity attacks
- Voice and multimodal inputs
- Deployment and CI/CD security
This broader approach helps security teams understand what an agent can access and how an attacker could move from one component to another.
222 Tests Help Measure Real-World Risk
The checklist assigns proposed severity levels to its tests:
- 75 Critical
- 108 High
- 30 Medium
- 9 Low
These ratings represent the potential importance of each test. They do not mean that every AI system contains the associated vulnerability.
Some tests focus on cloud and infrastructure weaknesses. Others examine how connected tools can be abused to move from an apparently harmless action to unauthorized data access.
Examples include:
- Stealing cloud credentials through SSRF
- Exploiting unsafe Python pickle loading
- Combining legitimate tools to transfer data without authorization
- Accessing another customer’s documents because of missing tenant controls
- Forging messages between AI agents
- Abusing delegated privileges
- Manipulating MCP-connected tools
Memory and retrieval systems receive particular attention. One test checks whether removing or bypassing a customer identifier could expose documents belonging to another tenant.
This is important because an AI application can appear secure while its underlying data-access controls remain weak.
Evidence and Safe Testing Matter
The checklist also provides guidance on how security teams should document their findings.
Tests can produce different types of evidence:
- Reflective evidence – The system directly shows the result.
- Blind evidence – The tester relies on changes such as timing or system state.
- Out-of-band evidence – A controlled callback confirms that an action occurred.
According to the checklist, reflective and controlled callback evidence can provide stronger confirmation, while blind-only results may need further validation.
The spreadsheet also includes testing objectives, recommended steps, suggested tools, expected results, severity ratings, and evidence requirements.
Security teams are advised to define the testing scope before starting, document excluded tests, and preserve supporting logs or screenshots. Destructive tests should only be performed when explicitly authorized, preferably in controlled or staging environments.
The checklist maps its tests to OWASP and MITRE ATLAS, giving security teams a structured way to connect individual tests with broader AI security frameworks.
The main takeaway is that agentic AI security cannot be limited to testing prompts. Organizations also need to examine the infrastructure, permissions, tools, memory, data flows, and connections that allow AI agents to take real-world actions.