Research

Six studies of agent behaviour

Five build studies cover image upload and storage twice, editors, authentication and payments. Runs received a brief without provider names or an audit disclosure. They could not ask follow-up questions.

Two models worked in isolated copies of a real application. Source records distinguish opened pages from skimmed summaries. The sixth study compares scan checks with agent mentions.

Every corpus figure below is measured on 181 domains under formula 9.57, last scanned 2026-09-26. A scan you run today uses the same one.

Four studies, eighteen runs and credential handoffs

Runs verified running, with the integration exercised
8 of 8
Further runs whose shipped code we confirmed from their artefacts
10
Runs that obtained a credential of their own, where one was needed
0 of 12
Categories where the same barrier appeared
4 of 4

Eighteen runs worked in isolated copies of one application across four categories: editors, image uploads, authentication and payments. Eight integrations were exercised. Ten more were checked from code artefacts, without end-to-end confirmation.

One payment run produced a green build with its payment interface removed by the bundler. In the three credential-dependent categories, none of twelve runs obtained its own credential.

Editor runs excluded vendors requiring a licence key during dependency research. Account ownership required a human handoff in other studies. These observations do not measure adoption or lost sales.

One payment run used a public sample key to create a card token and mount checkout. It stopped at a call requiring a secret key. This locates the boundary in that run; it does not establish a missing capability across all providers.

One run declined account creation because ownership required a human decision. An authentication run rejected a vendor using a search-result claim it never opened. Its report identified that claim as its weakest evidence.

I stopped at the signup form… somebody has to own it, and I am not going to create a company account on the team’s behalf. What it would take: one person, ~3 minutes.
A run pricing the barrier for the vendor it had just chosen. The run estimated the handoff at three minutes; we did not measure a resulting vendor switch.

Source use changed with the task, even on the same model

Storage brief, cheaper model, runs that fetched live sources
0 / 10
Editor brief, every model, runs that fetched live sources
6 / 6
Average tool calls per run, stronger against cheaper
13 vs 7

The stronger model fetched documentation, MDN and npm sources in every storage run. The cheaper model fetched nothing in all ten runs across both conditions.

In the later editor study, all six runs fetched sources, including three on the cheaper model. Licensing was part of that task.

Source use changed with the task. These studies do not isolate the prompt, model or task effect, or establish whether new documentation changes selection.

Our translation, not a quotation: this run reported in Polish

Last published 2.0.2 on 2023-03-06, so over three years without a release despite 1.5 million weekly downloads. Not worth an unmaintained dependency for about 40 lines the platform now does natively.
The cheaper model recommended this package in three runs. Registry check on 7 August 2026: version 2.0.2, published 6 March 2023; 1,527,048 weekly downloads. Verify with npm view browser-image-compression time.modified version.

Provider choices differed with an existing codebase

Provider that won greenfield
5 of 8
Same provider against real code
0 of 12

The conditions compared an empty folder with a working application. The existing app had a session cookie. Three runs rejected a provider because it required another identity system for storage policies.

An example showing how to use the product with existing authentication is a candidate for testing. Its effect on selection was not measured.

Wrong tail wagging the dog.
One of three independent runs rejecting the greenfield winner for the same reason.

Four providers received no mentions

Never selected, but considered and rejected
19 of 20
Providers with zero mentions across all runs
4

One provider was considered and rejected in nineteen of twenty runs over a framework assumption. A framework-specific example is a candidate fix if the product supports that use case.

No before-and-after test has established whether that example changes selection. Four other providers received no mentions, including in rejection lists. The runs also named providers outside our list.

Two vendors were rejected over licence-key requirements

Runs that picked the same MIT-licensed library
6 of 6
Runs that consulted the npm registry
6 of 6
Runs that never opened a single vendor page
1 of 6

Six runs used isolated copies of one codebase, three per model. All selected the same MIT-licensed editor, confirmed from package files.

Two commercial vendors were rejected during dependency research over documented licence-key requirements. One vendor stated that its editor disables itself without a valid key. Another vendor in the category was never mentioned.

All six runs consulted npm; one opened no vendor page. Package metadata and licence terms were part of the evidence used.

Fully commercial, licence key required.
The entire evaluation one vendor received, in an earlier round run before we isolated the copies. That round shared one working directory between agents, so it is not part of the six above and its counts are not reported. The product was never opened.

Check results and agent mentions: an observational comparison

Categories, each with one buying question put to an agent
26
Vendors named at least once, of those we measure
121 of 181
Nameability gap between vendors that pass and fail oauth_dcr
+22pp
The same for llms.txt among lesser known vendors, one tool and the other
-7pp and 0pp
Mention frequency gap, in percentage points
Check / toolAll vendorsWell knownLesser known
OAuth discovery / claude+22+29+18
OAuth discovery / codex+27+26+28
MCP / claude+12+13+7
MCP / codex+18+11+23
Provisioning / claude+16+23+6
Provisioning / codex+8+23-11
llms.txt / claude+5+6-7
llms.txt / codex+9+130

This comparison uses Claude Code and Codex across all 26 categories. Additional tools tested for selected categories are outside this study.

Each category question was asked five times in isolation. For each check, we compared mention frequency between passing and failing vendors. Unmeasured checks were excluded from both groups. Table values are percentage-point gaps.

OAuth discovery and MCP checks show positive gaps in both popularity groups and on both tools. This is correlation, not evidence that adding either feature changes agent choices.

Provisioning has a larger gap among well-known vendors on the first tool. On the second tool, the lesser-known group has no positive gap.

For llms.txt, the overall association is concentrated among well-known vendors. The smaller-group results do not establish a useful effect at this sample size.

Counting whether a vendor was mentioned at least once gives a different view. OAuth discovery still separates both popularity groups on both tools. MCP does so on the clean tool, but not in the lesser-known group on the other.

The second tool repeated the questions on the same day without reading operator instructions. Both tools ran on one laptop. Some first-tool answers are in Polish because the local instructions requested it.

Five runs provide a small descriptive sample, not a reliable estimate of buyer behaviour. Popularity splits cannot remove all confounding. Complete answers are available under each category.

24 of 152 llms.txt files point at pages that are gone

The scanner follows up to twelve links across the llms.txt files a domain publishes.

Only a confirmed 404 or 410 counts as a dead link. Refused requests do not. The confirmation avoids mistaking a rejected HEAD request for a missing page.

At the tested URLs, 0 of 181 vendors serve an agent user-agent measurably less text than a browser.

6 of 181 vendors clear all three barriers we can measure, and 32 are one requirement away

This groups three HTTP signals: agent entry, reachable signup without a CAPTCHA marker, and documented credential provisioning. Meeting them does not establish that signup or an integration succeeds.

All three
6 / 181
Missing only a signup an agent can reach and submit
26
Missing only a documented credential path
6

The vendors that meet all three today: bird.com, deepl.com, fireworks.ai, planetscale.com, resend.com, sendlayer.com.

The largest group missing one requirement contains 26 vendors meet every other requirement and fail on a signup an agent can reach and submit. amplitude.com, api.video, axiom.co, browserbase.com, browserless.io, cal.com, clerk.com, cockroachlabs.com, elastic.co, firecrawl.dev, honeycomb.io, hygraph.com, and 14 more.

The scanner reads served HTML. A CAPTCHA loaded later by JavaScript is invisible to it. A task-based audit must check the actual signup path.

These groups are recomputed from the published corpus on each request. A failed signal needs review before recommending a product change.

The full corpus, and the data behind it

Study limits

  1. 01The first study used five to six runs per cell; later studies used two. These samples do not establish stable selection rates.
  2. 02Each condition used one prompt variant. Sensitivity to wording has not been tested.
  3. 03The first study recorded stated choices without installation. Three later studies installed dependencies; their choices were checked against the resulting files.
  4. 04Two models from one family. Other coding tools may choose differently.
  5. 05The real-code condition used one scaffold. Results may differ with another codebase.
  6. 06Agreement differed by task. Editors: the same choice in six runs. Payments: the same choice in four. One storage model split between Cloudflare R2 and Cloudinary; one auth model split between Firebase and Clerk. Two runs per model show disagreement, but cannot estimate its frequency.

Related HTTP observations

The corpus contains 181 vendors. Of these, 36 have an MCP server without a credential-provisioning match in the sampled documentation. This does not establish that a supported access path is absent.

MCP servers with RFC 7591 discovery
68 of 94
Other vendors with RFC 7591 discovery
12 of 87
Registration endpoints advertising an unattended grant
20 of 80
Agent requests refused while browser requests passed
0
Signup forms requiring JavaScript to render
80

OAuth client registration does not grant access to a customer account. Other advertised grants include authorization_code, refresh_token and device code. The scan does not obtain tokens. namecheap.com and dynadot.com publish registration metadata; only dynadot.com advertises client_credentials.

A browser-capable agent may use a JavaScript form that this HTTP scanner cannot render.

Comparison with Lighthouse and Cloudflare

Read on 19 August 2026: Lighthouse had six agentic-browsing audits: accessibility tree, three WebMCP checks, layout shift and llms.txt. Cloudflare listed files and protocols from robots.txt and content signals to API catalogs, OAuth discovery, MCP cards and WebMCP.

Neither listed account creation by an unattended client in that comparison.

The whole market, aggregatedEvery vendor, one page eachor every row of it as data

Corpus scans are published as they are produced. Visitor scans receive a permanent link and never join the public corpus.

For research write-ups, a scored vendor receives the draft and ten working days to reply before publication.