Hy3
GPQA Diamond: 90.4%, HLE: 53.2%, BrowseComp: 84.2% (beats DeepSeek V4-Pro 83.4 & Opus 4.7 79.3; ~ties GPT-5.5 84.4), DeepSearchQA: 91.0, MCP-Atlas: 79.1 (beats Opus 4.7's 77.3), Terminal-Bench 2.1: 71.7% (beats Opus 4.7 & DeepSeek V4-Pro), SWE-bench Verified: 78.0%, GSM8K: 95.37%, MATH: 76.28%, Factual hallucination rate: 5.4% (down from 12.5% in preview), MRCR long-dialogue: 75.1% (up from 42.9%) | Grounded factual Q&A and RAG (trained to answer when grounded and flag missing evidence rather than fabricate), web-browsing/research agents, multi-step tool orchestration, customer-support and long multi-turn dialogue (coreference resolution, constraint tracking across turns), professional document generation, focused coding tasks.
Architectural Profile & Capabilities
GPQA Diamond: 90.4%, HLE: 53.2%, BrowseComp: 84.2% (beats DeepSeek V4-Pro 83.4 & Opus 4.7 79.3; ~ties GPT-5.5 84.4), DeepSearchQA: 91.0, MCP-Atlas: 79.1 (beats Opus 4.7's 77.3), Terminal-Bench 2.1: 71.7% (beats Opus 4.7 & DeepSeek V4-Pro), SWE-bench Verified: 78.0%, GSM8K: 95.37%, MATH: 76.28%, Factual hallucination rate: 5.4% (down from 12.5% in preview), MRCR long-dialogue: 75.1% (up from 42.9%) | Grounded factual Q&A and RAG (trained to answer when grounded and flag missing evidence rather than fabricate), web-browsing/research agents, multi-step tool orchestration, customer-support and long multi-turn dialogue (coreference resolution, constraint tracking across turns), professional document generation, focused coding tasks. Configurable reasoning effort: no-think default (low latency) plus low/high chain-of-thought for hard math, coding, and multi-step problems; supports interleaved thinking between tool calls. Tool-call stability generalizes across agent scaffolds (≤4% variance across CodeBuddy/Cline/KiloCode on SWE-bench Verified). Apache 2.0 licensed. **Caveats for ARMES: not for repository-scale refactoring (SWE-bench Pro: 57.9%, DeepSWE: 28.0% — use GLM-5.2 or DeepSeek V4 Pro), heavy numerical/statistical computation (defaults to naive allocations on quantitative optimization — pair with code execution), or contexts beyond its native 256K.**
Standard dense autoregressive transformer architecture with full-attention mechanisms.
Recommended Workloads & Primary Use Cases
Calls to Hy3 are routed through strict Zero Data Retention inference channels. Prompts and outputs are never stored, indexed, or monitored by Tencent.
More Models from Tencent
Compare other engines available in this laboratory.