All news
Research3 min

MARS8 Model Family: Technical Specifications and Benchmark Results

This technical report introduces MARS8 and presents detailed, open benchmarks evaluating its architecture, performance, and real-world TTS capabilities.

camb.ai/news
MARS8 Model Family

Introduction

Today, we introduce MARS8, our first full family of text-to-speech (TTS) models. Through real-world deployments across entertainment, sports broadcasting, and AI agents, we learned a fundamental lesson: no single TTS model can excel at every use case. Different domains demand different trade-offs between expressiveness, latency, controllability, and robustness.

Our journey began with MARS5, the first TTS system capable of handling high-intensity prosody at the level required for live sports commentary. MARS5 was open-sourced in English and quickly trended to #3 on Hugging Face, ranking just below models such as Gemini and Llama. We then expanded with MARS6 and MARS7, which became the first TTS models to be inducted into AWS Bedrock and Google Model Garden respectively.

In this document, we present the capabilities of MARS8, outline the different models within the family, and share detailed technical and academic benchmark evaluations. We evaluated MARS8-Pro and MARS8-Flash head-to-head against leading TTS systems in the industry, including Cartesia Sonic-3, ElevenLabs Multilingual v2/v3, and Minimax Speech-2.6-HD. All evaluations and datasets are fully open-source, ensuring that our benchmarks are transparent and replicable by the broader research and developer community.

What Makes MARS8 Different

Most TTS benchmarks evaluate models under ideal conditions: long, clean reference audio recorded in studio environments. But real-world applications rarely have this luxury. Users provide short clips, often recorded on phones, with background noise and natural expressiveness. We designed MARS8 to excel precisely where other models struggle. The defining capability is exceptional performance from limited data: even with audio as brief as 2 seconds, MARS8 maintains a highly consistent speaker identity and performance — something that typically requires 10-30 seconds of clean audio with competing systems.

MAMBA: The “Kobe Bryant” of TTS benchmarks

To validate these claims, we developed the MAMBA Benchmark, a rigorous stress test designed to reflect the most demanding real-world conditions rather than idealized studio environments. Three key design decisions make the dataset particularly challenging: 70% of samples require cross-language voice cloning; the average reference duration is just 2.3 seconds; and references contain natural expressiveness rather than neutral read speech. We are open-sourcing MAMBA so the broader community can independently replicate and validate our results.

Results

On speech quality, MARS8 leads on Content Enjoyment while matching the best performers on Production Quality. MARS8-Flash achieves a 5.67% Character Error Rate, demonstrating strong pronunciation accuracy across the multilingual test set. On speaker similarity — where MARS8 truly shines — MARS8-Pro achieves the highest scores on both metrics: 0.87 on wavlm-base-sv cosine similarity and 0.71 on CAM++. The CAM++ score of 0.71 represents a 38% improvement over the next best competitor and more than double the score of ElevenLabs Multilingual v2, all achieved with an average reference length of just 2.3 seconds.

The MARS8 Family

  • MARS-Flash (600M): ultra-low latency, TTFB as low as 100ms. For agentic conversations — call center and live conversation agents.
  • MARS-Pro (600M): balance of speed and fidelity, TTFB 800ms-2s. For expressive dubbing, audiobooks and digital media.
  • MARS-Instruct (1.2B): director-level emotion controls for high-end TV and film production, where speaker and prosody can be independently tuned.
  • MARS-Nano (50M): highly efficient for on-device applications, TTFB as low as 50ms. Deployed with partners like Broadcom.

MARS8-Flash, MARS8-Pro, and MARS8-Instruct are being released across multiple languages, collectively covering 99% of the world's speaking population, with Premium and Standard support tiers reflecting data scale (not output quality). Beyond this, CAMB supports up to 150 languages across the long tail.

Localize your content with CAMB.AI

Dub, translate, and voice your media in 150+ languages.

Get started for free