MARS8 is a family of production-grade multilingual text-to-speech models from Camb.ai, purpose-built for real-time and high-fidelity speech generation across 99% of the world's languages.
What is MARS8?
MARS8 is a set of four specialized TTS models that take text input and output natural-sounding speech with precise control over emotion, timing, and speaker identity. Deployable via API or directly on major compute platforms (Google Cloud, AWS, Hugging Face, Modal, Replicate, and more), it is designed by Camb.ai and powers live content like sports broadcasts and news where latency and accuracy are critical.
Key Features
- Four specialized models — MARS8-Flash (600M params, lowest latency for conversational AI), MARS8.1-Pro (600M, highest quality for dubbing/audiobooks), MARS8-Instruct (1.2B, director-level emotional control), and MARS8-Nano (50M, on-device inference).
- Global language coverage — Supports 99% of the world’s speaking population across premium languages (English, Hindi, French, Spanish, German, Japanese, Arabic, Korean, Chinese, Italian, Portuguese, Indonesian, Dutch, Russian, Tamil, Telugu, Bengali, Marathi, Kannada, Malayalam, Polish, Turkish) and regional variants (e.g., Arabic dialects, Brazilian Portuguese, Canadian French).
- Benchmark-leading quality — Outperforms Sonic-3, Speech-2.6-hd, and ElevenLabs Multilingual v2/v3 on PQ, speaker similarity (WavLM SV, CAM++), content enjoyment (CE), and word error rate (CER) per independent benchmarks.
- Production-grade economics — Optimized for low latency (sub-200ms TTFB for Flash), high concurrency, cost-efficiency at scale, and full privacy control via on-premises or private cloud deployment.
- Natively multi-cloud — First TTS model to launch simultaneously on all major compute platforms (Google Cloud, AWS, Hugging Face, Modal, Replicate, Baseten, Simplismart, Hathora, Parasail, Cerebrium, Fal.ai), eliminating API vendor lock-in.
- Fine-grained instruction control — MARS8-Instruct allows independent control of emotion, timing, and style separate from speaker identity, enabling director-level editing for film, TV, and audiobook production.
Who is it for?
- Live broadcasters — Use MARS8-Flash for real-time multilingual commentary in sports and news, where every second counts.
- Media and content producers — Dub films, shows, and audiobooks with MARS8.1-Pro or Instruct, achieving studio-quality output with expressive prosody and accent control.
- Conversational AI developers — Build voice agents for contact centers and live assistants using MARS8-Flash’s low latency.
- Embedded systems engineers — Deploy MARS8-Nano on automotive head units, edge devices, and IoT hardware where memory is constrained but voice quality must still exceed legacy TTS.
What can you do with MARS8?
- Real-time voice agents — Power customer support bots or virtual assistants that respond in 200ms or less across 20+ languages.
- Expressive dubbing — Dub a film into 40+ languages while preserving the original actor’s emotion and timing using MARS8-Instruct.
- On-device TTS — Run high-quality speech generation on a 50M-parameter model inside a car infotainment system or smart speaker without cloud connectivity.
How does MARS8 work?
You send text and optional voice reference (or emotion tags like [Excited], [Sad], [Laugh]) to the Camb.ai API or deploy the model weights directly on your chosen cloud platform. The model outputs a WAV or MP3 audio file with the specified speaker characteristics. For real-time use, MARS8-Flash streams audio after processing the first few tokens.