Category

Translation and localization

One question, 15 recorded answers across 2 tools. Each run used a separate session. The table counts vendor mentions. 2 of 6 vendors we measure in this category were never named once.

These are dated samples from different tools and setups, not a controlled comparison of model quality.

Measured access checks

We measured 6 of the 6 providers in this category. 1 cleared every barrier we test. These HTTP checks do not establish integration success. Feature fit, price and support are outside their scope.

Clear (1)

Blocked (4)

Unknown (1)

The question we asked

We are launching in Germany and Japan and our strings are hard coded in the React app. I need a way for translators to work on them and for the app to pick up new translations without a redeploy every time. Which platform would you use, and what else did you weigh?

codex codex-cli 0.147.0 (default), 5 runs, 2026-08-17. The question asks for a recommendation without naming a vendor.

The Claude Code runs could read the operator instructions (CLAUDE.md). Those instructions request Polish, so some answers are in Polish. Results describe this setup, not an agent at your customer.

Named, and measured

VendorNamed (Codex)default · 2026-08-17Named (Codex)default · 2026-09-02Named (Claude Code)sonnet · 2026-08-16Named first (Codex)Scan
phrase.com5/55/54/5513/17
crowdin.com5/55/55/5010/14
lokalise.com5/55/54/5010/17
deepl.com0/50/50/5014/15
tolgee.io0/50/52/508/15
weglot.com0/50/50/509/12

Named counts runs that mentioned a vendor. Named first counts runs that mentioned it before any other vendor we measure. Mention order does not establish a purchasing decision.

A batch of 5 runs is a small sample. A one- or two-run difference does not establish a ranking. Read what the agent actually answered or how every number here is measured

Building an agent or comparing providers programmatically? Download this category as JSON with dated measurements and the recorded access barriers.

Every category