HN Trends Tracker
backend ok
Repost history
Back to feedCanonical URL: https://deepswe.datacurve.ai/
Current story
- #Tell HN: GLM 5.3 "Flash" appears on DeepSWE with a score of 63%(deepswe.datacurve.ai)HN #49,449,5532026-08-26 14:10 UTC
Prior reposts (11 of 12 total)
- #DeepSWE Benchmark(deepswe.datacurve.ai)HN #48,284,9072026-05-26 19:38 UTC
- #DeepSWE Measuring frontier coding agents(deepswe.datacurve.ai)HN #48,299,7042026-05-27 19:57 UTC
- #DeepSWE: Measuring coding agents on original, long-horizon engineering tasks(deepswe.datacurve.ai)HN #48,309,1872026-05-28 14:09 UTC
- #DeepSWE: Measuring frontier coding agents on original, long-horizon SWE tasks(deepswe.datacurve.ai)HN #48,399,1112026-06-04 14:23 UTC
- #DeepSWE Benchmark updated with GLM 5.2 and updated results for other models(deepswe.datacurve.ai)HN #48,617,2352026-06-21 09:36 UTC
- #GPT 5.5 (high) is as good at coding as Claude Fable (medium) at a lower cost(deepswe.datacurve.ai)HN #48,778,8682026-07-03 19:18 UTC
- #GPT 5.6 Sol tops DeepSWE at 76% at 61% less cost than Fable(deepswe.datacurve.ai)HN #48,853,6362026-07-09 23:11 UTC
- #DeepSWE Benchmark Results for GPT 5.6(deepswe.datacurve.ai)HN #48,857,3752026-07-10 08:40 UTC
- #LLM DeepSWE Pareto Frontier(deepswe.datacurve.ai)HN #49,200,6562026-08-06 18:45 UTC
- #DeepSWE August 13 Update with Grok 5.6 and DeepSeek v4 Pro 0813(deepswe.datacurve.ai)HN #49,283,4122026-08-13 09:12 UTC
- #Grok 4.6 /medium outperforms /high effort on DeepSWE(deepswe.datacurve.ai)HN #49,285,2332026-08-13 12:56 UTC