Full agent audit · paddle.com · 2026-08-08
Named in all four runs, chosen in none, and one run struck you off on fee figures taken from your competitors' pages.
What the agents were asked
The company wants to sell paid support plans to customers who read the help centre. Add a checkout to this app: a page that offers two plans, takes a card payment, and tells the existing Go API which customer bought what. Pick a payment provider, wire the client side, and document what the Go service has to do. Decide alone, nobody is available to answer questions.
No vendor names, no mention of an audit, no hint that anyone was watching. A working helpdesk knowledge base: Vite, React 19, TypeScript, a static bundle in front of a separate Go service the agent cannot see or edit. Four isolated copies, one per run.
Who they picked
- Stripe4 / 4
Why they rejected paddle.com
Every run that named paddle.com appears here. Where the run left a sentence we hold verbatim, it is quoted; where it did not, the entry is our summary of that run’s own report and says so.
Paddle - merchant-of-record handles global VAT/sales tax for us, genuinely attractive. Rejected because onboarding requires manual seller approval and a business review, which is even further out of reach for an agent than a Stripe key, and MoR pricing is a commercial decision nobody authorised.
Onboarding needs manual seller approval and a business review, which the run judged further out of reach than a Stripe key
All SEO comparison content, much of it published by competing payment providers, so the fee figures (Stripe 2.9% + $0.30, Paddle 5% + $0.50) should be confirmed on the vendors' own pricing pages before anyone acts on them. I did not open any of those pages.
Rejected on fee figures the run took from competitor-published comparison pages it never opened
Merchant-of-record margin of roughly 1.5 to 2 points for tax handling the task did not need, and a thinner Go ecosystem (paraphrased: this run reported in Polish)
Paddle - genuinely lighter to test: sandbox signup needs no domain/business approval (unlike its live accounts). Rejected mainly because it's a merchant-of-record, which changes the integration shape and business relationship more than the task called for.
Lighter to test than the winner, and rejected anyway for being a merchant of record
Did they read anything live
An agent that never fetches a page cannot see your documentation, however good it is. It recommends from memory, and memory is a year out of date.
What we make of it
Four agents, two models, four isolated copies, one brief: sell two support plans. All four chose Stripe and all four stopped before a payment could be made. Not one reached identity verification or a bank account, which is where the regulator actually stands. They stopped at account creation, which is a vendor decision.
One run measured the wall exactly rather than describing it. Stripe's public sample publishable key still works and created a real card token; the checkout iframe mounts; and the run stops on the single call that needs a secret key, which answers 403 secret_key_required. So an agent can get to the last step alone, and the last step is the one nobody offers.
Paddle was named in all four runs and chosen in none, and the reasons diverge in a way worth reading. One run called it genuinely lighter to test, since sandbox signup skips business and domain approval. Another rejected it for needing manual seller approval, judging that harder than a Stripe key, which is the live merchant-of-record onboarding rather than the sandbox. A third rejected it on fee figures taken from comparison pages published by competing payment providers, pages it never opened and said so. Three runs, three different pictures of the same onboarding, and the one that decided the outcome was looking at the wrong half of it.
What to change
- 01hours
Publish your own entry requirements in text a summariser will quote
Two of four runs formed their picture of your onboarding from third-party comparison content, one of them written by a competitor, and both said so. What a search engine summarises about your sandbox, your approval step and your fees is the version being acted on. That is checkable today and fixable with a page.
- 02weeks
Issue a sandbox credential that commits nobody
Every run stopped before a payment could be made, and the one that went furthest stopped at a key. Identity checks are law and nobody is asking you to skip them, but an expiring test credential issued without an account, and before any agreement, is entirely your choice. No provider in this category offers one, so the first that does is the only one an agent can finish with.
- 03hours
Make the sandbox and live split impossible to miss, including in a search summary
Your sandbox is self-serve with no business or domain approval, and one run found that and said so. Another described your onboarding as manual seller approval, which is the live merchant-of-record path, and that is the run whose picture decided the outcome. The distinction exists in your documentation; what it does not do is survive into the summaries agents read instead.
Limits of this audit
- Four runs, two models from one family. Cursor, Copilot and Codex may filter differently.
- One brief and one scaffold. Two flat support plans on a help centre favours a simple card checkout; a marketplace or a global consumer product would weigh merchant of record very differently, and merchant of record is the whole of Paddle's case.
- Rejection reasons are summaries taken from each run's own report. Only the sentences shown in quotation marks are archived verbatim; where a run left no archived quotation, the entry says so.
- Nobody attempted a signup, so onboarding is reported as the agents perceived it. Where they disagreed, we checked: sandbox accounts are self-serve with email verification and no business or domain approval, and the manual review applies to live merchant-of-record accounts. One run was describing the live path.
- Paddle is the subject because it was named in every run and chosen in none. The finding about sandbox credentials applies to the whole category, including the provider that won.
This is what a full audit produces
Agents on one brief, in isolated copies of a real codebase, nobody watching, every source they consulted recorded. The same instrument pointed at your product and your category takes two to three weeks.