Techcrunch.com
Qwen 3.8-Max and Claude Opus 5 show why raw benchmark scores don't predict the bill
Alibaba released Qwen 3.8-Max this week and marketed the preview as second only to Claude Fable 5 (their launch-day table was more equivocal: the model leads on one of 12 coding-agent rows). But an independent harness came close to the opposite conclusion: a benchmark run, apparently using the Preview version, put Qwen 3.8-Max's best effort setting mid-pack, and its default setting last.Both results are real and defensible. The gap between them is about token and time budgets, and that matters...
Article source
Feeds.feedburner.com
All content rights belong to the original source
Related articles
Feeds.feedburner.com
How NTT DATA AIVista closes the last mile of agentic AI for enterprise agents
Feeds.feedburner.com
Slack wants to drag AI coding out of the terminal and into the group chat
Techcrunch.com