5 Comments
User's avatar
Scenarica's avatar

The two things Claude keeps winning have something in common: neither has a scoreboard. Reasoning and coding are benchmarked publicly, so effort concentrates there and the field converges. Writing and slide structure are judged on preference, which means no lab can prove it closed the gap and no buyer can prove it didn't. That's why your folder of side-by-side examples is doing work no evaluation suite currently does, and why the difference is likely to outlast the ones that are measured.

Stephen Smith's avatar

It's a great observation. I constantly test the models writing abilities (I guess I would lightly classify it as "taste"). They continually get better but I find that Claude just consistently does a better job for me in the type of writing that I like and it follows directions on writing exceptionally well. I also find it delivers a higher level of polish on powerpoint, word docs, and excel files too. Even though the latest version of ChatGPT can do a "good" job on building them - Claude is just better. Thank you for the comment - I appreciate it.

brian piercy's avatar

Give that good doggo some head scratches.

The Innovation Attorney's avatar

Do you find Fable is better at writing than the lower priced models?

Stephen Smith's avatar

Good question. I don’t actually. I’ve been extremely happy with Opus (since 4.5). I don’t find Sonnet to be a great writer. I use Fable for planning and Opus for execution. Glad you asked!