Things AI Is Genuinely Bad at That Nobody in Tech Wants to Admit

Things AI Is Genuinely Bad at That Nobody in Tech Wants to Admit

I spent eleven minutes last Tuesday watching ChatGPT confidently explain how to fix a plumbing issue that would have flooded my bathroom. The instructions were clear, well-formatted, and completely wrong about which valve to turn first. This wasn't a hallucination in the traditional sense — the AI wasn't making up facts. It just had no idea what "wrong" meant in a context where water pressure and pipe age actually matter.

The Confidence Problem Is Worse Than the Accuracy Problem

Everyone in AI circles talks about hallucinations. Models making up citations, inventing court cases, fabricating statistics. That's real, and it's a problem. But here's what two years of daily testing taught me that nobody mentions at conferences: the stuff AI gets subtly wrong is far more dangerous than the stuff it gets obviously wrong.

When ChatGPT invents a fake Supreme Court case, you catch it. When it gives you mostly-correct advice about negotiating a lease but misses the one clause that actually matters in your state — you don't catch it. You sign the lease. You find out six months later.

I tested this deliberately. I fed Claude and GPT-4 real scenarios from my own life where I already knew the right answer. Contract questions where I'd consulted an actual lawyer. Tax situations where I had the IRS response in hand. Medical questions where I had the diagnosis from an actual doctor.

The AIs were right about 70% of the time. That sounds pretty good until you realize you have no way of knowing which 70%. And they presented the wrong 30% with identical confidence. No hesitation. No "I'm not sure about this specific case." Just smooth, authoritative prose that happened to be incorrect.

Sequential Reasoning Falls Apart Fast

Here's a test I run on every new model: I describe a situation with three interlocking constraints and ask for a solution. Not math — just regular life logistics. Something like: "I need to be at Location A by 2pm, but I have a call at 1:30 that usually runs 20 minutes, and Location A is a 25-minute drive assuming no traffic, but today there's construction on the main route."

Humans immediately get it. You either reschedule the call, take it from the car, or accept you'll be late. The AI? It gives me a cheerful plan that somehow assumes I can teleport, or it confidently suggests leaving at 1:45 when the math obviously doesn't work.

I've done this test probably fifty times across different models. They fail it around 60% of the time. Not because the individual facts are wrong — they can tell me 1:30 plus 20 minutes is 1:50. They just can't hold multiple constraints in mind and trace through the implications.

This is the thing nobody in tech wants to say out loud: these models are not actually reasoning. They're pattern-matching against training data that contains reasoning. When the pattern is clear, they nail it. When it's not, they produce confident nonsense.

The Kick: What Happens When You Ask About Itself

The most revealing thing I've discovered — the thing that changed how I think about all of this — happened when I started asking AI models to predict their own failures.

I'd give ChatGPT a task, and before it answered, I'd ask: "What's the most likely way you'll get this wrong?" The responses were fascinating and completely useless. It would list generic AI limitations that had nothing to do with the specific task. Or it would claim limitations it doesn't have. Or — this happened more than once — it would say it might hallucinate, then immediately hallucinate in its actual answer without any self-awareness that it was doing the thing it just warned about.

The models have been trained to talk about AI limitations in the abstract. They'll recite the disclaimer language perfectly. But they cannot actually recognize when they're about to fail. There's no internal "I'm uncertain" signal that translates to output. The uncertainty language is just another pattern they've learned to produce when prompted.

I tested this by asking for confidence ratings alongside answers. The ratings had almost no correlation with actual accuracy. An answer rated 9/10 was wrong just as often as one rated 6/10. The confidence number was decoration.

Why the Tech Industry Won't Say This

Billions of dollars are riding on the premise that these models will get good at everything given enough scale. Admitting there are categories of tasks where they fail structurally — not just because they need more training data — threatens that narrative.

But after testing this stuff daily for two years, I'm pretty confident about a few structural limits. Anything requiring real-world feedback loops. Anything where context changes the meaning of "correct." Anything involving sequential constraints. Anything where the right answer depends on what you specifically care about rather than what's generally true.

I still use these tools constantly. They're genuinely useful for a lot of things. But I've stopped expecting them to know what they don't know. And I've noticed that the people most bullish on AI capabilities tend to be the ones who've tested them least carefully in their actual domain.

Maybe that's the uncomfortable question: what happens to an industry built on hype when the early adopters start comparing notes?

Heads up: Some links in this post may be affiliate links. I only recommend tools I've personally tested. Opinions are entirely my own.

댓글

이 블로그의 인기 게시물

How to Use Claude AI to Organize Your Messy Inbox (Without Losing Your Mind)

What Happens When You Give AI Tools a Real Deadline to Meet

I Handed My Entire Summer Trip to AI — Here's the Honest Breakdown After Two Weeks of Testing