Terminal-Bench 2.1: 86.6% (vendor) / 81.3% (AA independent), SWE-bench Pro: 67.7%, DeepSWE 1.1: 56.6%, PaperBench: 93.0% (#1), GPQA Diamond: 92.6%, FrontierSWE: 73.5%, OSWorld-Verified: 86.1%, OmniDocBench 1.5: 92.1%, IFBench: 82.8%, HLE: 43.6%, JobBench: 53.4% | Alibaba's most capable model and the first Qwen-Max released as open weights (GA Aug 3, 2026; open weights Aug 12, 2026).
Alibaba Foundation Models (1)
Available under ARMES unified subscription without separate API keys or billing agreements.
Tier:
Explore Other AI Labs on ARMES
Compare models across the 21+ leading laboratories unified in your workspace.