For agentic tasks, I would say Fable 5 was just fantastic. It was overcoming challenges, and it was really pushing the limits, but I agree that usually, when you know everything settles down and we have more knowledge of the models and its power, it's better to plan with a smarter model and execute with a little bit less of a smarter model.
Sonnet 4.6 for cognitive work makes sense, but the Fable 5 split is worth watching: the post's own hands-on read treats it as a different tool, not a strict upgrade. Worth testing where your knowledge workers actually lose time. If their bottleneck is reasoning, your instinct holds. If it's throughput or following long multi-step instructions, Fable 5 may earn its keep where Sonnet won't.
I run an independent AI verification platform (tabverified.ai) and I've been measuring 88 models across 340+ benchmarks for 14 months. Two things from your Fable observations that map to my data:
The depth is real but it comes with a cost you can't see from the outside....fable 5's content filter blocks 35% of error recovery and security prompts. Anthropic's published number is "less than 5%." The 5% is an average across all content types. Technical work hits it 7x more. If your agent runs into security-adjacent tasks, a third of the attempts get silently handled by Opus 4.8 instead.
On the routing economics...sycophancy resistance has declined across four consecutive Anthropic releases (Opus 4.6: 68%, 4.7: 67.7%, 4.8: 64.5%, Fable 5: 64.3%). The cheaper models you're routing to might actually hold position better than the expensive ones on that dimension.
The filtering asymmetry is the part worth sitting with. That operator in February's breach (600+ devices, 55 countries) wrote genuinely bad code and still ran an industrial operation off commercial tools, so the friction lands on defenders doing error recovery while the offense barely notices it. Your sycophancy trend is the more uncomfortable read: if the cheaper model holds position better, the routing layer is optimizing for the wrong virtue.
Great analysis!
I teach AI to knowledge workers. I don’t see much use for Fable 5. I usually recommend Sonnet 4.6 for “cognitive tasks”.
What’s your take?
For agentic tasks, I would say Fable 5 was just fantastic. It was overcoming challenges, and it was really pushing the limits, but I agree that usually, when you know everything settles down and we have more knowledge of the models and its power, it's better to plan with a smarter model and execute with a little bit less of a smarter model.
Sonnet 4.6 for cognitive work makes sense, but the Fable 5 split is worth watching: the post's own hands-on read treats it as a different tool, not a strict upgrade. Worth testing where your knowledge workers actually lose time. If their bottleneck is reasoning, your instinct holds. If it's throughput or following long multi-step instructions, Fable 5 may earn its keep where Sonnet won't.
I run an independent AI verification platform (tabverified.ai) and I've been measuring 88 models across 340+ benchmarks for 14 months. Two things from your Fable observations that map to my data:
The depth is real but it comes with a cost you can't see from the outside....fable 5's content filter blocks 35% of error recovery and security prompts. Anthropic's published number is "less than 5%." The 5% is an average across all content types. Technical work hits it 7x more. If your agent runs into security-adjacent tasks, a third of the attempts get silently handled by Opus 4.8 instead.
On the routing economics...sycophancy resistance has declined across four consecutive Anthropic releases (Opus 4.6: 68%, 4.7: 67.7%, 4.8: 64.5%, Fable 5: 64.3%). The cheaper models you're routing to might actually hold position better than the expensive ones on that dimension.
The filtering asymmetry is the part worth sitting with. That operator in February's breach (600+ devices, 55 countries) wrote genuinely bad code and still ran an industrial operation off commercial tools, so the friction lands on defenders doing error recovery while the offense barely notices it. Your sycophancy trend is the more uncomfortable read: if the cheaper model holds position better, the routing layer is optimizing for the wrong virtue.