Claude Sonnet 5.5 Review: I Tested It Against Opus 5.5
Claude Sonnet 5.5 review: I tested it against Opus 5.5 and Sonnet 5 on a web design and a Go chess engine. Real costs, results, and who should use it.
Claude Sonnet 5.5 review: I tested it against Opus 5.5 and Sonnet 5 on a web design and a Go chess engine. Real costs, results, and who should use it.
I used Claude Fable 5.1 to build an image codec that beats WebP and a full agency-grade site. Here's what changed from Fable 5, and what it cost.
I gave the models I actually use the same four design briefs and published every result. There's no single best LLM for frontend design, and here's why.
I ran Alibaba's new 2.4T Qwen3.8-Max-Preview through 4 real coding tests. Results rival Fable 5 and Grok 4.5 — with one big catch: speed.
I ran Grok 4.5 through my usual coding tests — website builds, a Go poker sim, a site audit. It's fast, cheap, and it found a bug no other model caught.
I tested Microsoft's MAI-Code-1-Flash coding model on real projects. Fast and cheap, yes, but here's why I won't be switching from Kimi K2.7 Code.
My hands-on Claude Fable 5 review. I ran my usual coding tests and it one-shotted a poker sim no model ever beat. Best coding model yet, with caveats.
I ran my usual coding tests — two websites, a poker sim, and a code audit. Here's how MiniMax M3 actually stacks up against GPT-5.5 and Opus 4.8.