Discussion about this post

User's avatar
Scenarica's avatar

The two things Claude keeps winning have something in common: neither has a scoreboard. Reasoning and coding are benchmarked publicly, so effort concentrates there and the field converges. Writing and slide structure are judged on preference, which means no lab can prove it closed the gap and no buyer can prove it didn't. That's why your folder of side-by-side examples is doing work no evaluation suite currently does, and why the difference is likely to outlast the ones that are measured.

brian piercy's avatar

Give that good doggo some head scratches.

3 more comments...

No posts

Ready for more?