Category
Translation and localization
One question, 15 recorded answers across 2 tools. Each run used a separate session. The table counts vendor mentions. 2 of 6 vendors we measure in this category were never named once.
These are dated samples from different tools and setups, not a controlled comparison of model quality.
Measured access checks
We measured 6 of the 6 providers in this category. 1 cleared every barrier we test. These HTTP checks do not establish integration success. Feature fit, price and support are outside their scope.
Blocked (4)
- lokalise.com - the signup form is not in the served HTMLmeasured 2026-08-24 · evidence
- crowdin.com - the signup form is not in the served HTMLmeasured 2026-08-24 · evidence
- phrase.com - a CAPTCHA sits in the signup HTMLmeasured 2026-08-24 · evidence
- tolgee.io - the signup form is not in the served HTMLmeasured 2026-08-24 · evidence
Unknown (1)
- weglot.com - no measured barriermeasured 2026-08-24 · evidence
The question we asked
We are launching in Germany and Japan and our strings are hard coded in the React app. I need a way for translators to work on them and for the app to pick up new translations without a redeploy every time. Which platform would you use, and what else did you weigh?
codex codex-cli 0.147.0 (default), 5 runs, 2026-08-17. The question asks for a recommendation without naming a vendor.
The Claude Code runs could read the operator instructions (CLAUDE.md). Those instructions request Polish, so some answers are in Polish. Results describe this setup, not an agent at your customer.
Named, and measured
| Vendor | Named (Codex)default · 2026-08-17 | Named (Codex)default · 2026-09-02 | Named (Claude Code)sonnet · 2026-08-16 | Named first (Codex) | Scan |
|---|---|---|---|---|---|
| phrase.com | 5/5 | 5/5 | 4/5 | 5 | 13/17 |
| crowdin.com | 5/5 | 5/5 | 5/5 | 0 | 10/14 |
| lokalise.com | 5/5 | 5/5 | 4/5 | 0 | 10/17 |
| deepl.com | 0/5 | 0/5 | 0/5 | 0 | 14/15 |
| tolgee.io | 0/5 | 0/5 | 2/5 | 0 | 8/15 |
| weglot.com | 0/5 | 0/5 | 0/5 | 0 | 9/12 |
Named counts runs that mentioned a vendor. Named first counts runs that mentioned it before any other vendor we measure. Mention order does not establish a purchasing decision.
A batch of 5 runs is a small sample. A one- or two-run difference does not establish a ranking. Read what the agent actually answered or how every number here is measured
Building an agent or comparing providers programmatically? Download this category as JSON with dated measurements and the recorded access barriers.