Benchmark across Claude, OpenCode, Hermes, and other coding agents on 10 SWE-bench tasks

Carlo Capocasa published a 10-task GLM benchmark comparing Claude, OpenCode, Pi, Zcode, Hermes, and 3code on representative SWE-bench verified tasks. 3code solved 9 of 10 tasks using 5 million tokens, while Pi solved 6 of 10 using fewer tokens than other runners-up. Capocasa noted harness performance varies with token efficiency and task completion rates, cautioning that users should validate results against their own heuristics.

Topics

Agentic AIClaude

Sources

Go deeper

This intelligence is sourced automatically from public sources across the web and synthesised by the Prefactor AI pipeline. Stories are reviewed before publication.