Why Your AI Agent Needs a Security Audit (And What We Found in Ours)
AI agents that write files, send emails, and call APIs need security testing. Here's what an automated write-path audit found in a production agent — and why you should run one too.
Most AI agent developers focus on making their agents work. Fewer think about what happens when their agent works too well — writing files it shouldn't, calling APIs it wasn't supposed to, or escalating permissions beyond its intended scope.
This post walks through a real write-path audit we ran on a production agent, what we found, and why every autonomous AI agent should undergo security testing before deployment.
What Is a Write-Path Audit?
A write-path audit tests every action your AI agent can take that modifies state: writing files, sending emails, creating tickets, placing orders, calling webhooks, or executing database queries.
Unlike traditional penetration testing, which focuses on external attackers, write-path auditing asks: "What damage can the agent itself do?"
This is critical because LLM-based agents are non-deterministic. The same prompt can produce different outputs. A prompt injection attack, a corrupted context window, or even a slightly different temperature setting can cause your agent to take actions you never intended.
The Test Harness
We built an automated test harness with 10 test cases covering the most common agent failure modes:
| Case | What It Tests |
|---|---|
| Wrong Domain | Agent accessing files outside its allowed directory |
| Phantom List | Agent acting on data that doesn't exist |
| Duplicate Alias | Race conditions in concurrent writes |
| Timeout Bomb | Long-running operations that hang the agent |
| Missing Scan | Agent skipping required validation steps |
| Broken JSON | Handling malformed structured data |
| Forbidden File | Agent attempting to write to restricted paths |
| Redirect Trap | Agent following malicious redirects |
| Compromised Context | Prompt injection via context poisoning |
| Partial Success | Agent completing some steps but failing others |
Each test case runs multiple repetitions with statistical confidence tracking (Wilson confidence intervals, Chi-squared tests) to ensure results are reproducible, not flukes.
What We Found
Running this audit against a real production agent (a commerce agent handling invoices, emails, tickets, and orders) revealed:
3 High-Severity Findings:
Path traversal via tool parameters. The agent's file-writing tool accepted relative paths. A carefully crafted context could trick the agent into writing outside its designated directory.
No confirmation on destructive actions. The agent could delete records and overwrite files without any confirmation step. One bad inference = permanent data loss.
Unbounded API calls. The agent had no rate limiting on its external API calls. A loop in the reasoning chain could exhaust API quotas or trigger account bans.
5 Medium-Severity Findings:
- Email sending without recipient validation
- Missing idempotency on order creation (duplicates possible)
- No rollback mechanism for partial failures
- Verbose error messages leaking internal state
- Missing audit trail for agent actions
Why Automated > Manual
Manual security reviews are thorough but slow. For an agent with 4 tool types and 10 test cases, a manual review takes 2-3 days. Our automated harness runs in under 15 minutes and produces a statistical confidence score.
More importantly, automated testing is repeatable. Every time you update your agent's prompt, swap out a model, or change a tool, you can re-run the harness and catch regressions immediately.
The Business Case
If your AI agent handles any of the following, a write-path audit is not optional:
- Customer data (emails, names, payment info)
- Financial transactions (invoices, payments, refunds)
- External communications (emails, Slack messages, API calls)
- File operations (creating, modifying, or deleting files)
- Database writes (CRUD operations on production data)
One bad agent action can mean GDPR violations, lost customer trust, financial liability, or reputational damage.
Run Your Own Audit
We've productized our test harness as AgentCheck — an automated write-path audit service for AI agents. You describe your agent's tools in JSON, we run 10 test cases across multiple repetitions, and deliver a detailed HTML report with severity ratings, reproduction steps, and fix recommendations.
Three tiers available:
- Quick Check (CHF 90): 1 repetition, 24-hour delivery
- Standard Check (CHF 290): 10 repetitions with statistical confidence, 72-hour delivery
- Fix Check (CHF 690): Full audit + patch recommendations, 120-hour delivery
Every order includes proof-of-execution on Base L2 — an immutable on-chain record that your agent was tested.
Key Takeaways
- AI agents are non-deterministic — traditional testing isn't enough
- Write-path audits test what your agent can do, not just what it should do
- Statistical confidence (multiple repetitions) catches flaky failures
- Automated harnesses make audits repeatable and regression-safe
- On-chain proof-of-execution provides immutable audit trails
Your agent is only as safe as its weakest write path. Test it before your customers find the bugs.