Neural AI Voice & Playback
Listen to long-form completions and documents with 12 ultra-natural neural voices, pitch control, and scrubbing controls.
High-Fidelity Audio Synthesis
Reading dense research reports, multi-page strategy memos, or lengthy code reviews on a screen isn't always practical when you are commuting, exercising, or multitasking.
ARMES includes a built-in Neural Text-to-Speech (TTS) engine that converts any assistant response into studio-grade audio with natural pauses, emotional cadence, and clear pronunciation.
12 Distinct Neural Voices
Choose from 12 distinct neural voices powered by advanced speech synthesis models, each calibrated for distinct listening scenarios:
| Voice Name | Character & Tone | Ideal Context |
|---|---|---|
| Leda (Default) | Youthful, clear, and bright | Everyday answers, modern productivity, fast reviews |
| Kore | Firm, authoritative, and confident | Strategic memos, executive decisions, legal analysis |
| Puck | Upbeat, energetic, and lively | Creative brainstorming, product ideation |
| Zephyr | Warm, encouraging, and bright | Coaching, personal planning, continuous listening |
| Charon | Informative, steady, and measured | Complex technical documentation, research papers |
| Fenrir | Dynamic, bold, and excitable | High-impact announcements, key takeaways |
| Aoede | Breezy, natural, and conversational | Long-form articles, casual dialogues |
| Orus | Grounded, firm, and resonant | Financial analyses, governance reviews |
| Callirrhoe | Easy-going, smooth, and relaxed | Relaxed listening, narrative fiction, reflections |
| Autonoe | Bright, articulate, and crisp | Step-by-step tutorials, educational guides |
| Enceladus | Soft, contemplative, and breathy | Philosophical inquiries, introspective notes |
| Umbriel | Calm, steady, and easy-going | Evening reading, deep focus sessions |
Plan Availability & Privacy
Neural Voice playback is available on Pro and Ultra plans. All text-to-speech requests execute under account-level Zero Data Retention: audio is synthesized on-demand in memory and never logged or used for voice training.
Playing Audio in the Interface

- Trigger Playback: Click the Speaker icon located in the action bar above or below any completed assistant response.
- Interactive Playback Bar: A persistent playback bar docks at the bottom of the workspace:
- Play / Pause / Restart: Immediate playback control with keyboard spacebar support.
- Interactive Timeline: Scrub forwards or backwards through the audio stream.
- Speed Multipliers: Toggle playback speeds from
0.75x,1.0x,1.25x,1.5x, to2.0x. - Audition in Settings: Test and customize your default voice in Settings → Voice.
Voice Input vs. Voice Playback (iOS App)
It is important to distinguish voice playback from voice input:
- Voice Playback (TTS): Converts assistant written responses into high-fidelity speech inside the web and iOS apps.
- Voice Input (Apple Siri & App Intents): On the ARMES native iOS app, you can use Apple Siri to speak inputs directly to your agents or capture notes hands-free:
- "Hey Siri, ask ARMES."
- "Hey Siri, create a note in ARMES titled Q4 Goals."
- Siri passes the dictated text securely into ARMES without exposing your workspace history to third parties.