SWE-bench Pro: 65.7%, GPQA Diamond: 92.3%, Terminal-Bench 2.1: 85.4%, DeepSWE v1.1: 64.3%, HLE: 55.4%, SWE-bench Verified: ~82.9%, NL2Repo: 58.9%.
Tencent Foundation Models (2)
Available under ARMES unified subscription without separate API keys or billing agreements.
Tier:
Explore Other AI Labs on ARMES
Compare models across the 21+ leading laboratories unified in your workspace.