Q: How do I secure AI agents?
An agent can act, and anything that can act can do damage. Security here means limiting what damage is possible, because you can't fully control what a model decides.
The core principle is least privilege. Give the agent only the tools its job requires, and make each tool as narrow as possible: read-only where read-only works, one folder instead of the whole disk, a spending cap where money is involved. Before adding any tool, ask the blast-radius question: if the model calls this with the worst possible arguments, what happens? If the answer scares you, narrow the tool or add a confirmation step.
For consequential actions (deleting data, sending emails, spending money), keep a human in the loop: the agent proposes, a person approves. For example, that's exactly how the Anthropic connector to Gmail works: an agent can read and draft emails, but it cannot delete or send them.
And treat everything the agent reads (web pages, documents, emails) as untrusted input, because attackers hide instructions in content exactly where agents will read them (this is called prompt injection).
MCP adds one more surface: a third-party MCP server is code you're trusting, like any dependency. Use servers from sources you trust, read what tools they actually expose, and prefer running them locally where you can see them.
Then watch the traces. An agent's behavior drifts with model updates and new inputs, and the trace log is where you catch it early.