Discussion about this post

User's avatar
Fabrice Talbot's avatar

Great analysis!

I teach AI to knowledge workers. I don’t see much use for Fable 5. I usually recommend Sonnet 4.6 for “cognitive tasks”.

What’s your take?

Keeping TABs on Your AI Agents's avatar

I run an independent AI verification platform (tabverified.ai) and I've been measuring 88 models across 340+ benchmarks for 14 months. Two things from your Fable observations that map to my data:

The depth is real but it comes with a cost you can't see from the outside....fable 5's content filter blocks 35% of error recovery and security prompts. Anthropic's published number is "less than 5%." The 5% is an average across all content types. Technical work hits it 7x more. If your agent runs into security-adjacent tasks, a third of the attempts get silently handled by Opus 4.8 instead.

On the routing economics...sycophancy resistance has declined across four consecutive Anthropic releases (Opus 4.6: 68%, 4.7: 67.7%, 4.8: 64.5%, Fable 5: 64.3%). The cheaper models you're routing to might actually hold position better than the expensive ones on that dimension.

3 more comments...

No posts

Ready for more?