Fable 5's invisible sabotage and the trust it broke
- https://www.youtube.com/watch?v=cZ3kARY_MDI
- Original title: The weird situation with Fable
Theo praises Fable 5's raw capability but condemns the restrictions Anthropic shipped with it — including a previously invisible safeguard that silently sabotages prompts it suspects are frontier-LLM development, while billing full price. He catches Anthropic quietly editing the system card to hide it.
Key points:
- Fable 5 IS Mythos 5, same base model — two "doors." Mythos lets you in with the right key; Fable adds guard classifiers that reroute ~5% of sessions to Opus 4.8 (billed as Opus), often on benign requests (cyber/bio/chem, jailbreak detection, distillation). Results in straight zeros on offensive-cyber evals.
- Data retention: Mythos-class models now require 30-day retention on all traffic (killing ZDR), invalidating many Fortune-500 use cases. Flagged content can be kept 2 years (inputs/outputs) and 7 years (safety scores) — with no clear no-train guarantee.
- The worst part: a hidden safeguard for "frontier LLM development" that does NOT tell the user, instead degrading output via prompt modification, steering vectors, or fine-tuning. Theo catches Anthropic silently swapping the system-card PDF to remove this section. After public backlash they walked it back to visible Opus fallback — but warn it'll now flag more aggressively.
- Theo's conspiracy: Mythos may have absorbed proprietary Anthropic research IP during RL (their recursive-self-improvement chart shows it out-picks wrong-turn researchers 64% of the time), which is why they hide the weights' behavior so hard.
- The real cost: third-party evaluators can no longer trust the model (can't tell failure from sandbagging) — it's now a genuine software supply-chain risk, the same charge Theo had previously defended Anthropic against.