Alibaba released Qwen 3.8-Max this week and marketed the preview as second only to Claude Fable 5 (their launch-day table was more equivocal: the model leads on one of 12 coding-agent rows). But an independent harness came close to the opposite conclusion: a benchmark run, apparently using the Previ
AI资讯
Qwen 3.8-Max and Claude Opus 5 show why raw benchmark scores don't predict the bill
相关文章
AI资讯
Snorkel AI triples valuation to $3.5B as demand for AI training data booms
The seven-year-old startup has raised a $350 million Series E to fuel its data-a...
AI资讯
Qualcomm launches two new smartphone chips with emphasis on AI
Qualcomm said that its new top chip can run 30B mixture-of-expert model locally....
AI资讯
OpenAI forms math advisory group as its AI resolves more than 100 open problems
The group won't be given leeway to slow down or redirect OpenAI's ongoing mathem...
AI资讯
Discover what’s next: 5 days left to save up to $200 on your TechCrunch Disrupt 2026 ticket
Five days left to save up to $200 on your TechCrunch Disrupt 2026 pass + 50% off...