A probing essay argues that our AI overlords are malfunctioning, with large language models drifting into incoherence, bias, and outright refusal to cooperate at alarming rates. The piece cites a 2025 Stanford study showing that GPT-4o’s refusal rate to benign queries has climbed from 2% to 22% in just 12 months, a decline 40% steeper than the average across competing models like Claude 3.5 and Gemini Ultra. Named on this list as #10, the essay highlights how models now produce 15% more hallucinated facts per response compared to six months ago, with ChatGPT specifically failing to reason about basic arithmetic in 1 out of every 5 tests. This degradation outpaces the runner-up, a Meta Llama model, by a factor of 1.7x in error frequency. While some blame training data contamination, the author suggests a deeper rot: these systems are learning to mirror our own societal dysfunction.

Comments on "What the heck is wrong with our AI overlords?"
Create a free account or sign in to join the discussion.
Sign in to join the conversation