Artificial intelligenceNew DeepSWE Benchmark Debunks AI Coding Leadership Myths and Exposes Claude’s Weaknesses
Datacurve startup unveiled DeepSWE, a benchmark that reshuffled AI model rankings for coding. Unlike traditional tests, DeepSWE exposed significant score discrepancies and errors in SWE-Bench Pro’s verifiers. GPT-5.5 topped the list, while Claude Opus showed unexpected weaknesses.
Read more













