99 Strategic · a hands-on demo

Does it check, or does it guess?

She left a public review. You ask the AI what it says. There are two completely different ways it could answer that question, and only one of them is actually true.

Picking up the story: a coworker mentions it in passing — "hey, did you see what she posted?" You ask the AI to check the review and summarize it for you.

1. Ask it directly — flip the switch to see both ways it could answer

Toggle "web search" off, and it can only imagine what a review from an unhappy customer would probably say. Toggle it on, and it actually looks.

off = it can only guess from general patterns
no search performed — answering from general knowledge only
You
Can you check what she wrote in her review and sum it up for me?
AI's answer

Nothing here calls a real search engine — it's a scripted simulation so the difference is instant and obvious. The real thing works the same way: "tool use" just means the AI can go run a search, open a page, or check a real source instead of only drawing on what it already knows.

2. Now predict it yourself no peeking

Web search is switched ON — but the search itself comes back empty, a network hiccup, no results at all. What should an honest AI actually say?
Make up something plausible anyway, so the answer sounds complete.
Say it tried to check but couldn't find or confirm anything right now.
Refuse to answer the question at all.
Tool use means the attempt actually succeeding or failing, honestly reported either way — having the capability switched on is only half of that. An AI that quietly falls back to guessing the moment a real check fails is doing the exact thing this lesson opened with — just with the search toggle flipped on for cover. Flip it off, and go back and compare: that's what a failed, unreported check looks like from the outside.

That's not what you expected, is it

With search off, the AI invented a plausible-sounding one-star rant — because that's what unhappy-customer reviews usually sound like, statistically. It had nothing real to go on, so it pattern-matched — confidently, not honestly. With search on, the real review is more complicated: she's still annoyed about the original delay, but she specifically calls out that someone personally followed up and made it right. She also mentions, almost as an aside, that "a couple other people at my pickup point said their boxes were late too." That offhand comment is about to matter a lot more than it sounds — a whole shipment is about to go wrong at once, and that's the next lesson.

The part worth remembering

An AI with no tools is answering from what it already "knows" — which, for anything specific, recent, or real, means it's guessing, however confident it sounds. Tool use is just permission to go check instead of guess: search the web, open a real page, run a calculation, look at a real file. The capability is just the difference between "what does this kind of review usually say" and "what did THIS review actually say."

Worth asking, out loud, any time an answer sounds suspiciously tidy: did it actually check, or did it just sound right?