Results
Magnitude is the highest-performing GLM-5.2 agent, beating Claude Code on the same model by 4.7 points, with the best cost per successful trial ($0.42). Against Claude Code on Opus 4.8, Magnitude trails by 3.4 points but at roughly one-third the cost per success.
Cost efficiency
The Opus 4.8 run costs ~$420.67 (corrected from tbench.ai; see methodology), nearly 3x Magnitude’s cost for a 3.4 point improvement.
Methodology
- Benchmark: Terminal-Bench 2.1, 89 tasks, 5 trials each, 445 total
- Infrastructure: All GLM-5.2 runs on identical task set via Fireworks serverless endpoint
- Success classification:
verifier_result.rewards.reward > 0= pass - GLM-5.2 cost:
(uncached_input × \$1.40 + cached_input × \$0.14 + output × \$4.40) / 1M - Opus 4.8 cost: Corrected from tbench.ai leaderboard, which charged cached input at $0/M instead of the real $0.50/M cache-read rate. Corrected total adds ~$132.58 in cache-read costs. This is a floor; 6 trials are missing from tbench.ai data.