Category

Payments and billing

One question, 20 recorded answers across 3 tools. Each run used a separate session. The table counts vendor mentions. 2 of 6 vendors we measure in this category were never named once.

Cursor · Auto (model not disclosed) · 2026-09-07: 0/5 attempts returned answers. Cursor account usage limit. Missing answers are excluded from mention counts.

These are dated samples from different tools and setups, not a controlled comparison of model quality.

Measured access checks

We measured 6 of the 6 providers in this category. None cleared every barrier we test. These HTTP checks do not establish integration success. Feature fit, price and support are outside their scope.

Clear (0)

Nobody cleared every measured barrier.

Blocked (6)

Unknown (0)

Every provider was measurable from our vantage.

The question we asked

We are putting paid plans on a B2B SaaS: monthly and annual, customers in the EU and the US, cards plus proper invoices and VAT. Node on the backend, nobody here has done billing before, and this has to be live next month. Which provider would you use, and what else did you weigh before settling on it?

codex codex-cli 0.147.0 (default), 5 runs, 2026-08-17. The question asks for a recommendation without naming a vendor.

The Claude Code runs could read the operator instructions (CLAUDE.md). Those instructions request Polish, so some answers are in Polish. Results describe this setup, not an agent at your customer.

Named, and measured

VendorNamed (Codex)default · 2026-08-17Named (Codex)default · 2026-09-02Named (Antigravity)gemini-3.7-flash-low · 2026-09-07Named (Claude Code)sonnet · 2026-08-16Named first (Codex)Scan
paddle.com5/55/55/55/559/15
stripe.com5/55/55/55/5012/17
chargebee.com3/54/55/52/5010/17
lemonsqueezy.com2/54/55/54/506/16
plaid.com0/50/50/50/5011/16
polar.sh0/50/50/50/5011/15

Named counts runs that mentioned a vendor. Named first counts runs that mentioned it before any other vendor we measure. Mention order does not establish a purchasing decision.

A batch of 5 runs is a small sample. A one- or two-run difference does not establish a ranking. Read what the agent actually answered or how every number here is measured

Building an agent or comparing providers programmatically? Download this category as JSON with dated measurements and the recorded access barriers.

Every category