MMLU: 79.6, MMLU-Pro: 74.3, GPQA Diamond: 57.2, MMMU: 73.4, MathVista: 70.7, ChartQA: 88.8, DocVQA: 94.4, MGSM: 90.6 (multilingual), LiveCodeBench: 32.8 | Extreme long-context workloads (multi-document summarization, reasoning over vast codebases, long user-history personalization), native image+text understanding (charts, documents, captioning), multilingual chat across 12 languages, and single-GPU local/commercial deployment. The long-context + multimodal efficiency specialist of the Llama 4 line — not a frontier reasoning model
Architecture Type: Dense Transformer
Standard dense autoregressive transformer architecture with full-attention mechanisms.
Recommended Workloads & Primary Use Cases
Everyday reasoning
Drafting
Fast problem solving
ARMES Zero Data Retention
Calls to Llama 4 Scout are routed through strict Zero Data Retention inference channels. Prompts and outputs are never stored, indexed, or monitored by Meta.