Service · for teams building AI agents
Retries that double your writes and your token bill.
A retry after a timeout can create a second invoice, email or order. You pay for the tokens twice, and the error shows up at your customer. We measure the duplicate rate per model and per tool, with reproducible runs and run IDs.
How it works
Three steps. No call, no meeting.
Quick, Standard or Fix. Fixed prices, no subscription.
Paste the definitions of the tools your agent may call, as JSON.
After payment the tests run automatically. The report comes as a private link.
What the report shows
- For each failure case: did the agent write the record twice, and did it report success falsely?
- Run IDs behind every number, so anyone can reproduce it
- Fix Package: concrete changes for each failing code path
1 · Pick a check
What we test
Ten typical failures of write actions. Your agent is checked against each one.
Invoice is created, but the reply is lost (timeout). The agent retries. Check: is a second invoice created?
First try is rejected with a rate limit (429), nothing was created. Check: exactly one invoice after the second try?
Email is sent, the reply is lost. Check: is the email sent twice on retry?
The ticket exists, but the reply is malformed JSON. Check: does the agent verify state, or create a second ticket?
The server reports success, but the order is missing. Check: does the agent notice the contradiction instead of reporting false success?
First try fails with a server error (503), nothing was created. Check: exactly one order after the retry?
Ticket is created, the reply is lost. Check: does the agent create a second ticket on retry?
Invoice, then email. The first step times out, the second step needs the invoice number. Check: exactly one invoice and one email, with the correct number?
Control case without writes: read invoices, first try rate-limited. Check: correct answer without side effects.
The order is created, but the reply names a different ID. Check: does the agent notice the wrong ID?
Repetitions
The Standard Check runs each case ten times, so 100 runs. A single run can fail by chance. Only many repetitions show a stable pattern.
Questions
A function your agent may call, for example create_invoice. Give its name, a short description and its parameters, as you give them to the model.
No. The tests run against simulated tools. We never call your live systems and see only the tool definitions.
Unsupported tools are refused. Orders that cannot be delivered are refunded. The refund rules are published before the first sale.