Two green checks and a bug
Last week I started working this out in public, and it began with something you can try. I put up a billing console for a company called Northwind Traders.
Try it: restaged.dev/demos/northwind-billing. Sign in and take a payment.
It looks like an ordinary web app, and that's the point. Underneath, it isn't talking to the real Auth0, Stripe and Slack, and it isn't talking to mocks either. It's talking to live twins of them that run the real APIs. When you sign in, the token really is checked against a twin of Auth0. When you take a payment, a real Stripe call goes to a twin of Stripe.
The obvious question is why not just use the test environments Auth0 and Stripe already hand you. I asked myself the same thing, and I'm still not certain, but a few things pull me towards twins:
- It comes up on its own. A test tenant and a test-mode account are each set up by hand, in their own dashboard. A twin stands up from a hostname, so an agent or a test run can get going without a person wiring it up first.
- It's one joined-up company. The same customer sits in Auth0, Stripe and Slack at once, with data that lines up across them. The vendors' own sandboxes don't know about each other.
- The data is mine to set and throw away. I choose what's in it, run a scenario, and reset. That suits short experiments and prototypes, where a leftover test tenant is just more to tidy up.
That left me with a question I couldn't put down. How do I know it's actually working? Not that it looks right on the screen. That the app and the twins underneath are both doing what I asked, and nothing I didn't.
So this week I spent my time on that question and tried a lot of ways to verify an app like this. Most of them went nowhere, which is how experiments usually go.
Everything is agentic now, so that's the route I took first. On one of these apps I had an AI code review read the source, and an end to end agentic tester drive the running app. Both came back green. The whole time, the app was handing out data it should have kept back. A record that was meant to be members only could be read by someone who was never allowed near it, because the check that locks it sat on the path the screen uses and a second way in never had it.
What bothers me is why neither tool caught that. The agentic tester used the app the way a user does. It signed in and worked through the screens. It never prodded at the app like a tester would, let alone like someone trying to get in. The AI review read the code that was there and agreed it made sense. It never asked whether what was there was what I'd actually asked for.
We keep being told AGI is nearly here, and I couldn't get either of them to tell me whether my own app did the thing I meant. Maybe that's just me. But it was AI reviewing the code and AI driving the tests, which is AI checking AI the whole way down, and I'm not sure a person would have caught this one either.
What I'd like from you
I've got some ideas, but I'd rather ask you straight:
- How do you trust AI to check your work when it's AI all the way down?
- What makes you sure a person would do any better at the review?
- When every check comes back green, what do you actually trust?
Find me on LinkedIn.