Category
Transactional email APIs: who an agent names, and who it never mentions
We put one question to an agent 5 times on claude, every run a separate session with nothing carried between them, and counted the vendors it named. 2 of 6 vendors we measure in this category were never named once.
The question we asked
Our password reset and receipt emails go out through a box we run ourselves and too many of them land in spam. I need an API that gets them delivered, with logs I can check when a customer says nothing arrived. Node, maybe fifty thousand emails a month. Which provider would you use, and what else did you look at first?
claude 2.1.233 (Claude Code) (sonnet), 5 runs, 2026-08-16. The question names no vendor and asks for a recommendation, which is the shape a developer types.
Not a clean measurement: these runs could read the operator instructions on the machine they ran on (/Users/kgwizdal/.claude/CLAUDE.md), which is also why some answers are in Polish rather than English: those instructions ask for it. They describe an agent there rather than an agent at your customer, and we say so rather than publish the number alone.
Named, and measured
| Vendor | Named (claude) | Named first | Scan |
|---|---|---|---|
| postmarkapp.com | 5/5 | 4 | 10/15 |
| sendgrid.com | 5/5 | 0 | 9/14 |
| resend.com | 4/5 | 1 | 14/16 |
| mailgun.com | 2/5 | 0 | 9/14 |
| loops.so | 0/5 | 0 | 13/16 |
| sendlayer.com | 0/5 | 0 | 14/16 |
The two columns answer different questions. Named is whether you were in the room at all. Named first is whether you were the answer. A vendor at zero is not losing on price or features in these runs: it is not being considered.
5 runs separate a wall from silence and nothing finer. Two vendors a run or two apart are not ranked by this, and we would rather say that than sell the gap. Read what the agent actually answered or how every number here is measured