She left a public review. You ask the AI what it says. There are two completely different ways it could answer that question, and only one of them is actually true.
Toggle "web search" off, and it can only imagine what a review from an unhappy customer would probably say. Toggle it on, and it actually looks.
Nothing here calls a real search engine — it's a scripted simulation so the difference is instant and obvious. The real thing works the same way: "tool use" just means the AI can go run a search, open a page, or check a real source instead of only drawing on what it already knows.
With search off, the AI invented a plausible-sounding one-star rant — because that's what unhappy-customer reviews usually sound like, statistically. It had nothing real to go on, so it pattern-matched — confidently, not honestly. With search on, the real review is more complicated: she's still annoyed about the original delay, but she specifically calls out that someone personally followed up and made it right. She also mentions, almost as an aside, that "a couple other people at my pickup point said their boxes were late too." That offhand comment is about to matter a lot more than it sounds — a whole shipment is about to go wrong at once, and that's the next lesson.
An AI with no tools is answering from what it already "knows" — which, for anything specific, recent, or real, means it's guessing, however confident it sounds. Tool use is just permission to go check instead of guess: search the web, open a real page, run a calculation, look at a real file. The capability is just the difference between "what does this kind of review usually say" and "what did THIS review actually say."
Worth asking, out loud, any time an answer sounds suspiciously tidy: did it actually check, or did it just sound right?