I build most of the ranch's software with an AI agent now. That's not the hard part. The hard part is that the same agent tells me how it went, and for a long time I had no way to check.

So I started handing its work to the competition.

Anything sensitive here goes only to providers I trust — that was Claude and my local models for a while. I added Grok on July 27 and OpenAI on the 31st. Once all three were inside that line, I could send one model another's raw code and ask what's wrong with it.

Last week my agent compared two coding models, got three out of three on both, and wrote down "non-inferior." I'd have believed that. GPT read the same code and said the test proved nothing — both sides scored perfectly, so neither had room to fail. My own notes had made that point two days earlier.

That's what I didn't see coming. I thought I was buying second opinions. What I got was a way to check a claim before I acted on it, including my own claims. Every real change gets one now. It costs pennies.

Then I ran this post through it, and both of them killed my favorite line.

I'd written that when two rivals agree on something, it's almost certainly true. Wrong, they said, for the same reason: they're trained on much the same material, so agreement isn't two independent opinions. It's a reason to go look, not proof.

I stopped reading my agent's summary of its own work. I read what its rivals said about it instead — including here.