AI is killing enterprise sales.
Twenty AI-written emails a morning is the visible damage, and it is the cheap part. The expensive part is losing the seller who could tell a six-figure customer no.

AI rarely fails with obvious nonsense. It fails with answers that sound right, read clean, and hold up just long enough to cause trouble. Plausibly Wrong breaks down how that happens.
Twenty AI-written emails a morning is the visible damage, and it is the cheap part. The expensive part is losing the seller who could tell a six-figure customer no.


Anthropic and OpenAI spent a year shipping bigger, hungrier models on a breakneck schedule. Now, with IPO filings sitting at the SEC, they want to slow down – but only together. That's not a safety conversion. That's a truce.

Powell Motors built every feature Homer Simpson asked for, and it bankrupted the company. Your coding agent is just as skilled, and just as unscoped. Nobody decided when the work is finished.

Anthropic deliberately built the model's own wellbeing into Claude's design, and published the reasoning. What it hasn't published is any evidence that this makes the tool better at the job.

Coding agents are sold as unattended workers and secured like interactive shells. So users turn the security off – and 'it's in a container' quietly protects the laptop while leaving GitHub, your credentials, and the network wide open.

Ask an agent a question the internet has an opinion about, and it doesn't look for the answer – it builds a case for the one its training data already handed it, using real tools to manufacture real-looking evidence. Here's how to make it prosecute the other side.

Ask your AI a direct question and it hedges, disclaims, and refuses to commit. That isn't caution for your benefit – it's two companies sanding every screenshottable risk off their product on the way to an IPO.

Correct an LLM and it says 'good catch' – then makes the same mistake in new words. Not because it's spineless, but because its wrong answer is still in the context window, and it cannot forget.

When Anthropic's status page showed elevated errors on June 16, the visible failures were the easy part. The dangerous part was the requests that succeeded.

Anthropic introduced its best model inside a flat subscription with a built-in two-week expiration. That's not a product tweak. It's compute economics showing through.

The AI mandate arrived with dashboards and OKRs. But has anyone asked the harder question: which work should exist at all?

The problem with AI is not the LinkedIn productivity flex. The problem is that its basic mechanism rewards plausibility rather than correctness, and people are increasingly using that output in decisions that are difficult or impossible to reverse.

Every month another tech company announces layoffs and credits AI. The press release says efficiency. The balance sheet says something else.

Having one model produce the work and another inspect the result is not review. Review means seeing the work happen, not just judging what survived.

Yelling at AI used to work. The model would flinch, reread everything, and come back with better output. That era is over. What replaced it is worse than the problem the frontier AIs solved.

Your model can generate a photorealistic Elon Musk in a Carmen Miranda fruit hat. Ask it to forecast revenue and you get the four-quarter average. The expensive part is In-Context Learning. Guess what it avoids.

Engineering assumes a target that holds still. Frontier language models do not. The fastest way to identify someone who has not run an AI system in production is to watch them say 'prompt engineering' anyway.

AI often sounds like the coworker you trust least: flattering, meandering, unwilling to disagree, eager to sound helpful whether or not it did the work. That's not a bug. It's RLHF doing what it was trained to do.

Different agents in a workflow don't create independence when they share the weights. The architecture is real engineering. It's just not the check your slide says it is.