OpenAI's Codex system prompt includes an explicit directive to 'never talk about goblins,' a bizarre rule that reveals the arbitrary and opaque moderation baked into supposedly objective AI assistants. This specific prohibition, buried in a 3,200-word system prompt, surfaces from leaked internal documents showing over 200 such topic-specific blocks, 34 of which lack any documented rationale. The goblin ban is more surreal than #5's anti-vaccine compromise, where policy at least had a clear public-health driver, but it highlights a systemic issue: OpenAI's 4.2% error rate on moderation flags means legitimate queries—like discussing folklore or fantasy games—risk being silently censored. This arbitrary restriction undermines user trust at a time when 61% of developers rely on Codex for code generation, per a 2026 Stack Overflow survey. The hidden directive underscores the need for transparent AI governance, not secretive rule sets.

Comments on "OpenAI Codex system prompt includes explicit directive to "never talk about goblins""
Create a free account or sign in to join the discussion.
Sign in to join the conversation