We caught our own AI naming the wrong deal from four numbers in plain sight — twice. Here's what it taught us about DIY diligence with ChatGPT-style tools.
This week we were doing something unglamorous: re-testing our own AI on a question every searcher asks in their first five minutes — "Which of my deals has the highest asking price?"
Four deals. Four numbers. One right answer, sitting in plain sight.
It got it wrong. So we tightened the instructions and asked again — and it got it wrong a second time, this time with complete confidence: the headline named one deal while its own list showed a bigger number two lines below.
Here's the uncomfortable part: this wasn't a bug, and it wasn't a bad model. It was a frontier AI model — the same class of model behind every general-purpose chatbot people paste CIMs into — doing what these models quietly do. Ask one to find several numbers scattered across a pile of documents and compare them while writing a fluent answer, and it will sometimes anchor on the wrong thing and deliver the mistake beautifully. Polish and correctness are independent. You cannot instruct your way out of it. We tried — all we got was a wrong answer with better posture.
Why this matters to you: every searcher running diligence through a raw chatbot is running this exact experiment on a real deal, with real money, and no second ask to catch the miss.
Our fix wasn't a better prompt. We took the job away from the model: where an answer is computable — a ranking, a comparison, a figure in your own records — it's now computed deterministically from the data before the AI says a word. The model's job is language; arithmetic belongs to machinery that can't be charming and wrong at the same time. Re-asked the same question: right answer, first try, and right for a structural reason. A thousand asks, the same answer a thousand times.
We caught this on internal test data before a single user hit it — because we run the same "verify, don't trust" discipline on our own product that we preach for deal documents.
Curious how others here are handling AI in their diligence workflow — are you double-checking its arithmetic, or trusting the fluency?
(We're tearing down a real listing live on Sept 9 — claimed vs. supported vs. missing, line by line. Event's posted here on Searchfunder if you want to watch the discipline applied to an actual deal.)