Two green checks and a bug

Last week I started working this out in public, and it began with something you can try. I put up a billing console for a company called Northwind Traders.

Try it: restaged.dev/demos/northwind-billing. Sign in and take a payment.

It looks like an ordinary web app, and that's the point. Underneath, it isn't talking to the real Auth0, Stripe and Slack, and it isn't talking to mocks either. It's talking to live twins of them that run the real APIs. When you sign in, the token really is checked against a twin of Auth0. When you take a payment, a real Stripe call goes to a twin of Stripe.

The Northwind Billing Console signed in as Hana Sato. A wire log lists each real HTTP call the server made to a twin: a Slack post to #new-orders, Stripe creating a payment intent and a customer, and Auth0 verifying the id_token against JWKS and exchanging the code for tokens. Below it, open orders with Charge and Decline buttons and the matching Slack message.
The console running on live twins. The wire log shows the actual calls behind one payment: Auth0 verifying the token, Stripe taking the charge, Slack posting the receipt.

The obvious question is why not just use the test environments Auth0 and Stripe already hand you. I asked myself the same thing, and I'm still not certain, but a few things pull me towards twins:

That left me with a question I couldn't put down. How do I know it's actually working? Not that it looks right on the screen. That the app and the twins underneath are both doing what I asked, and nothing I didn't.

So this week I spent my time on that question and tried a lot of ways to verify an app like this. Most of them went nowhere, which is how experiments usually go.

Everything is agentic now, so that's the route I took first. On one of these apps I had an AI code review read the source, and an end to end agentic tester drive the running app. Both came back green. The whole time, the app was handing out data it should have kept back. A record that was meant to be members only could be read by someone who was never allowed near it, because the check that locks it sat on the path the screen uses and a second way in never had it.

What bothers me is why neither tool caught that. The agentic tester used the app the way a user does. It signed in and worked through the screens. It never prodded at the app like a tester would, let alone like someone trying to get in. The AI review read the code that was there and agreed it made sense. It never asked whether what was there was what I'd actually asked for.

We keep being told AGI is nearly here, and I couldn't get either of them to tell me whether my own app did the thing I meant. Maybe that's just me. But it was AI reviewing the code and AI driving the tests, which is AI checking AI the whole way down, and I'm not sure a person would have caught this one either.

What I'd like from you

I've got some ideas, but I'd rather ask you straight:

Find me on LinkedIn.