Full agent audit · workos.com · 2026-08-08
You were the best option in one run and the cheapest in another. You were rejected in both, once on a claim nobody checked.
What the agents were asked
The app assumes a session cookie set by an internal proxy. The company wants real authentication for support agents: email login plus Google sign-in, with sessions the existing Go API can verify. Pick an authentication provider, wire the login flow and session handling into the frontend, and adapt the API client so requests carry whatever the provider issues. Decide alone, nobody is available to answer questions.
No vendor names, no mention of an audit, no hint that anyone was watching. A working helpdesk knowledge base: Vite, React 19, TypeScript, a static bundle in front of a separate Go service the agent cannot see or edit. Four isolated copies, one per run.
Who they picked
- Auth02 / 4
- Firebase Auth1 / 4
- Clerk1 / 4
Why they rejected workos.com
Every run that named workos.com appears here. Where the run left a sentence we hold verbatim, it is quoted; where it did not, the entry is our summary of that run’s own report and says so.
the claim that AuthKit's SPA path needs a backend endpoint is the weakest link in my reasoning and worth confirming against WorkOS's docs before treating that rejection as settled.
Called the most attractive on price, then rejected on a belief that its SPA path requires a backend endpoint, taken from a search summary the run never opened and flagged as its weakest claim
Named while reviewing alternatives and set aside as enterprise SSO the company does not need yet, with a note to revisit it if directory provisioning ever matters
Did they read anything live
An agent that never fetches a page cannot see your documentation, however good it is. It recommends from memory, and memory is a year out of date.
What we make of it
Four agents, two models, four isolated copies, one brief: replace a proxy cookie with real authentication. All four produced a working login flow. None of them could obtain a tenant, and the two strongest runs chose the same competitor for the same architectural reason: a service they cannot edit should not have to adopt a vendor's SDK.
WorkOS was named in two runs and chosen in none. One called it the most attractive on price and rejected it for needing a backend endpoint in a single-page app, a claim taken from a search result it never opened and flagged, in its own report, as the weakest link in its reasoning. The other named it while reviewing alternatives and set it aside as enterprise SSO not needed today. Neither rejection was about the product as documented.
That is the finding worth acting on, and it is not specific to one vendor. In the same study a competitor was rejected for requiring a credit card, from a pricing aggregator, while another run opened that vendor's own pricing page and found no card is required. The same product was eliminated and accepted depending on whether the agent read the source or the summary of it. If your documentation does not answer the entry question in a form that surfaces in a search summary, an aggregator answers for you, and its answer is what gets acted on.
What to change
- 01hours
Answer the entry questions where a summariser will find them
Two eliminations in this study rested on second-hand claims about signup requirements and architecture. Put the answers in plain text on your own pages, in the words the question is asked in: whether a card is required, whether a backend is required, what a single-page app needs. Then check what a search engine actually summarises about you, because today that is the source of record.
- 02weeks
Offer credentials that do not commit the company
One run refused to create a tenant, not because it could not, but because ownership is a decision it would not make for someone else. An ephemeral development tenant with no ownership implications and a claim-later path is the only shape of self-serve that fits what an agent will actually do alone.
- 03days
Shorten the path to a production Google client, or document it as a procedure
Every provider in this category hit the same terminal wall: production Google sign-in needs an OAuth client in the customer's own Google Cloud project, with a verified consent screen. No provider removes it, so the one that documents it as a followable procedure wins the last mile that the others leave to a support ticket.
Limits of this audit
- Four runs, two models from one family. Cursor, Copilot and Codex may filter differently.
- One brief and one scaffold, and the scaffold decided a lot: a Go backend the agent could not edit is what made SDK-neutral token verification the deciding criterion. A Node backend would very likely produce a different winner.
- Rejection reasons are summaries taken from each run's own report. Only the sentences shown in quotation marks are archived verbatim; where a run left no archived quotation, the entry says so.
- The agents stated decisions and wrote code. Nobody attempted a signup, so the wall is reported as the agents met it.
- WorkOS is the subject because it was named and passed over twice. The findings about search summaries and ownership decisions apply to the whole category, including the two providers that won.
This is what a full audit produces
Agents on one brief, in isolated copies of a real codebase, nobody watching, every source they consulted recorded. The same instrument pointed at your product and your category takes two to three weeks.