Sesame: CSM 1B
Sesame: CSM 1B is a neural speech synthesis model by Sesame, providing expressive voice rendering for real-time conversation.
Run Sesame: CSM 1B
Included on PRO plan
Context Window
128K
128K tokens
Input Token Price
$7.00
per 1,000,000 tokens
Output Token Price
$7.00
per 1,000,000 tokens
Speed Rating
Fast
Fast
Architectural Profile & Capabilities
Sesame: CSM 1B powers ARMES Voice mode, translating conversational text streams into low-latency, natural audio without storing audio logs.
Architecture Type: Neural Audio Synthesis
Standard dense autoregressive transformer architecture with full-attention mechanisms.
Recommended Workloads & Primary Use Cases
Conversational voice mode
Hands-free listening
Audio playback
Multi-lingual speech
Operational Caveats & Boundaries
- •Per-character or per-token voice synthesis rates apply.
ARMES Zero Data Retention
Calls to Sesame: CSM 1B are routed through strict Zero Data Retention inference channels. Prompts and outputs are never stored, indexed, or monitored by Sesame.