← Blog

Introducing Crosstalk 1 and Crosstalk 1 mini

14 July 2026

Crosstalk 1 and Crosstalk 1 mini are our new text-to-speech (TTS) models, optimized for mobile hardware. They stream speech faster than real time and start speaking almost instantly. They can also clone voices from short samples, emit word timings as they go, and have a compact download and memory footprint.

Crosstalk 1

A high-fidelity model for iOS.

Platforms
iOS
Compatibility
iPhone 15 up
Download
iOS: 534 MB
Peak memory
iOS: ~1.5 GB

Crosstalk 1 mini

Smaller and faster cross-platform model.

Platforms
iOS, Android
Compatibility
iPhone 12 Pro up, Fairphone 4 (2021) up
Download
iOS: 131 MB Android: 146 MB
Peak memory
iOS: ~2 GB Android: ~1 GB

Both are built on open-weight checkpoints that we've optimized and extended to give the best possible user experience.


Comparisons

Below we compare the Crosstalk models against other TTS models in their weight class. Models with cloning support were given the reference voice below to copy:

The reference voice clip.

TTFA (time to first audio) is how long you wait before the first sound of speech. A low TTFA makes generation feel instant.

RTF (real-time factor) is how long generation takes divided by the length of the audio it produces. Anything below 1 is faster than playback, and lower is always better – it widens the range of hardware that can keep up and lightens thermal load.

We tuned the Crosstalk models heavily for low TTFA and RTF to offer a near-instant, seamless speech generation experience.

Crosstalk 1

End-to-end time taken to generate the ~12s clips below
15.9 s
9.1 s
8.4 s
2.9 s
NeuTTS Air 748MMarvis TTS 250MQwen3-TTS 0.6BCrosstalk 1
Crosstalk 1 against its weight class, measured on an iPhone 17 Pro
ModelDownloadPeak memoryWord timingsCloningStreams
Crosstalk 1460 ms0.24534 MB1.3 GBliveyesyes
NeuTTS Air 748M0.9 s1.30839 MB1 GBnoyesyes
Qwen3-TTS 0.6B1.4 s0.722.0 GB4.5 GBnoyesyes
Marvis TTS 250M1.4 s0.68940 MB1.9 GBnoyesyes

Crosstalk 1 mini

End-to-end time taken to generate the ~12s clips below
8.1 s
6.2 s
5.6 s
3.0 s
2.2 s
0.6 s
MOSS-TTS-Nano 100MNeuTTS NanoKokoro 82MSupertonic 3 99MCrosstalk 1 miniPiper 16M
Crosstalk 1 mini vs smaller models, measured on a Samsung Galaxy S25+
ModelDownloadPeak memoryWord timingsCloningStreams
Crosstalk 1 mini85 ms0.19146 MB870 MBliveyesyes
NeuTTS Nano610 ms0.56506 MB1.3 GBnoyesyes
MOSS-TTS-Nano 100M9 s0.60717 MB6.8 GBnoyesyes
Supertonic 3 99M3.7 s0.22129 MB720 MBnonono
Kokoro 82M2.1 s0.47349 MB1.2 GBnonoyes
Piper 16M210 ms0.0567 MB910 MBnonoyes

On-device benchmarks

The comparisons above use one flagship phone per table; this section shows how the Crosstalk models perform across a wider range of iOS and Android devices.

Crosstalk 1

DeviceYearChip
iPhone 12 Pro2020A141500 ms1.1
iPhone 152023A16900 ms0.54
iPhone 16 Pro2024A18 Pro700 ms0.53
iPhone 17 Pro2025A19 Pro460 ms0.24

Crosstalk 1 mini

DeviceYearChip
iPhone 12 Pro2020A14110 ms0.71
iPhone 152023A1670 ms0.50
iPhone 16 Pro2024A18 Pro80 ms0.42
iPhone 17 Pro2025A19 Pro50 ms0.26
Fairphone 42021Snapdragon 750G350 ms0.78
Pixel 82023Tensor G3210 ms0.48
Galaxy S25+2025Snapdragon 8 Elite85 ms0.19

Details

Hardware acceleration

We optimized Crosstalk 1 for the Apple Neural Engine, which is why it runs best on recent iPhones. Crosstalk 1 mini has to run on Android too, so it uses the CPU and GPU on iOS, and the CPU alone on Android.

Word timings

Subtitles need word timings, and the usual way to get them is to run speech-to-text over the finished audio, or to force-align the script against it.

The Crosstalk models emit the timings themselves, so we get word-aligned subtitles for free. They're accurate enough to drive live, karaoke-style highlighting as the speech is generated, with no speech-to-text and no aligner in the pipeline.

Production-ready

Alongside driving RTF and TTFA down, we also cut the download size and peak memory so that the models fit into an app and run on as wide a range of hardware as possible.

In practice, Crosstalk 1 mini is small enough to ship inside an app bundle, so there's no separate model download before you can start generating speech. Crosstalk 1 is about a third smaller than the next comparable model, and both are far smaller and require less memory than the open-weight checkpoints they started from.


Crosstalk 1 mini ships in our flagship app later this month.