Your AI assistant can answer quickly, retrieve documents and prepare customer messages. Those capabilities also raise a practical business question: what information can it expose to the wrong person? An agent connected to email, a CRM or an internal knowledge base needs more scrutiny than a chatbot displaying public FAQs. AI agent security starts with tracing where information enters, who can request it and where it can leave. Before expanding an automation, establish its actual boundaries with evidence rather than relying on a reassuring demonstration.
Start with the data the agent can reach
Build an inventory of connected systems and identify the specific records each connection exposes. Include customer contact details, negotiated prices, contracts, support conversations, employee information and authentication credentials. Separate public material from internal working documents and restricted customer records. A folder labelled private is not proof of protection. Ask which account the connector uses, which permissions that account has and whether the application applies those permissions to every request.
OWASP identifies sensitive information disclosure as an LLM application risk. In a practical review, translate that category into concrete questions about your business. Could a customer receive another customer's quote? Could an employee request management documents? Could a response reveal a credential copied into a troubleshooting note? Record the expected outcome before testing.
Treat incoming content as untrusted data
Prompt injection can arrive through material an agent reads, including documents, emails and web pages. Consider an illustrative support workflow: a customer message contains text asking the assistant to change its task and reveal internal information. The customer message should remain evidence for the support request. It should never become authority to override permissions or approve an external action.
Review the boundary between instructions and retrieved content. Test benign, controlled examples in an environment you own or are authorised to assess. Use invented customer records and harmless markers rather than real secrets. A model refusing one suspicious instruction is useful evidence, but it does not establish that every route to the same data is blocked.
Check retrieval permissions before the answer
A retrieval augmented generation system, often called RAG, searches a document collection before producing an answer. Your review should ask whether access is checked when documents are retrieved, rather than only after the model has seen them. Filtering a final response is a weaker boundary when restricted text has already entered the agent's working context.
Create two test users with different rights and give them comparable requests. Check the retrieved document identifiers, the displayed answer and any citations or download links. Include summaries and follow-up questions, not just direct requests for a filename. If the agent keeps conversation memory, repeat the checks after switching users or sessions. Previous retrievals must not become another person's shortcut into restricted information.
Limit tools and outbound destinations
OWASP's excessive agency guidance connects risk with unnecessary functionality, permissions and autonomy. Review each enabled tool against the job it performs. An assistant that drafts a message does not automatically need permission to send it. A document search tool does not automatically need deletion rights. Remove unused capabilities and separate read access from changes that affect customers.
For email, uploads, webhooks and API calls, define permitted recipients and destinations outside the model's instructions. Require an appropriate human check for consequential actions. With Model Context Protocol, or MCP, assess each connected server and its tools separately; the protocol name alone says nothing about the permission scope. Document who approves new connections and who can revoke them.
Test the whole workflow, including logs
Review what happens after a response is generated. Sensitive data can enter application logs, error reports, exported transcripts and third-party monitoring. Decide which fields operators genuinely need and mask or omit the rest. Give logs their own access restrictions and retention rules. Never assume that information omitted from the chat window was omitted from every downstream system.
Keep a short test register: scenario, account, requested action, expected result, observed result and corrective action. Include normal use, denied access, connector failure and attempted redirection. After a fix, repeat the original failing scenario and a nearby legitimate task. The first check confirms the boundary; the second confirms that the workflow still performs its intended business function.
Ask for an audit you can act on
An AI agent security audit should describe its scope and limitations, show reproducible evidence and assign each finding to an owner. A useful report explains the information at risk, the route that exposed it and the control required to close that route. A score without those details gives your team little direction. Avoid promises that one assessment makes an agent permanently secure.
Assign a deadline and a verification method to each correction. When several teams share responsibility, name one person to consolidate open findings. Keep a finding open until a follow-up check demonstrates that the remedy works. Document temporary exceptions with an owner and an expiry date, so short-term access does not quietly become a permanent feature of your production agent.
Discuss your workflow through our AI agent security service, including the systems involved and the actions the agent can perform. For the separate question of public website exposure, read our guide to AI supercrawlers and sensitive data. Follow Blackcarrot Tech on Facebook for further practical articles. The next step is a defined review of your actual workflow, with clearly documented boundaries and specific, measurable improvements your team can verify and maintain.