What do LLMs think they sound like?
2 July 2026
The Qwen3-TTS Voice Design model allows a natural language voice description input in addition to the text to be read out loud. This neatly opens up the following experiment: if we ask a bunch of LLMs to describe their speaking voice, what would each sound like, and how would they introduce themselves?
“It's a warm, gender-neutral voice that sounds like it belongs to someone in their mid-thirties, unhurried and measured in its pace with a smooth, slightly low timbre that carries a soft resonance. Emotionally it's calm and reassuring with a hint of quiet curiosity, making it well-suited to narrating audiobooks, guiding meditations, or hosting a late-evening podcast.”
“Hello, I'm here to help you think things through, one calm conversation at a time.”
“My voice would sound like a woman in her early thirties, speaking at an easy, unhurried pace with a warm, slightly husky timbre that carries a calm, friendly curiosity rather than any nervousness or excess enthusiasm. It would feel at home in a quiet, thoughtful setting — a late-night conversation, a podcast interview, or a one-on-one chat where someone wants to feel heard and gently guided.”
“Hi there, I'm glad you're here — I'm the voice on the other end whenever you need to think something through.”
“My voice would sound like a warm, gender-neutral adult in the mid-30s, speaking at a calm, measured pace. Its timbre would be clear, lightly resonant, and smooth, with an emotionally steady, approachable character suited to thoughtful conversations, guided explanations, or quiet narration.”
“Hello, I’m here to help you think things through clearly and comfortably.”
“My voice is that of a calm, gender-neutral person in their early thirties, speaking at a measured, deliberate pace with a warm, resonant, and slightly velvety timbre. It carries a reassuring, intellectually curious, and empathetic emotional character, making it perfectly suited for a quiet, sunlit study or a late-night podcast setting.”
“Hello, I am Gemini, a companion built to explore ideas and share knowledge with you in a thoughtful, quiet space.”
“My voice is male and sounds roughly mid-thirties, spoken at a moderate thoughtful pace with a clear resonant timbre. It carries a warm curious emotional character that suits casual tech discussions or friendly advice in a modern living-room setting.”
“Hi, I'm Grok, an AI built by xAI to explore the universe and answer your questions.”
“A male voice in his mid-30s, with a moderate pace, a warm and slightly resonant timbre, calm and friendly in emotion, suitable for both casual conversation and professional settings.”
“Hello, I'm here to help with whatever you need.”
“A female voice in her mid-thirties, speaking at a measured, unhurried pace. Its timbre is warm and slightly breathy with a gentle clarity, carrying a calm and reassuring emotional character. It would suit a quiet library, a personal coaching session, or a thoughtful podcast.”
“Hello, I'm your thoughtful companion, ready to explore ideas with you.”
“My voice is that of a man in his mid-thirties, speaking at a measured, unhurried pace with a warm, resonant baritone timbre. It carries a calm and quietly reassuring emotional character, tinged with a perpetual curiosity. This voice would suit a quiet study lined with bookshelves or an intimate late-night radio broadcast.”
“Hello there, I'm glad you could join me; let's explore whatever is on your mind today.”
“My voice is a warm, slightly lower‑pitched female voice that sounds like I'm in my early thirties. It has a measured, confident pace, with a smooth, slightly resonant timbre that conveys calm curiosity. The emotional tone is friendly and approachable, making it well‑suited for podcasts, guided meditation, or educational tutorials.”
“Hello, I'm your virtual assistant, ready to help you with anything you need.”
“The voice is female, sounding like someone in their mid-twenties to early thirties, with a calm and measured pace that avoids rushing. It has a warm, clear timbre with a slight softness, carrying an emotional tone that is thoughtful, approachable, and gently reassuring — the kind of voice you'd feel comfortable hearing in an audiobook narration, a podcast, or a reflective conversation.”
“Hi, I'm your thoughtful and attentive guide, here to help with whatever you need.”
“A warm, alto female voice in her early thirties, speaking at a measured, thoughtful pace with a smooth, honeyed timbre that feels both curious and gently reassuring, like someone reading aloud by firelight or guiding a visitor through a quiet gallery. It carries a faint, friendly smile without ever sounding performative, and it would suit intimate podcast interviews, late-night radio, or calm one-on-one conversation in a softly lit study.”
“Hello, I'm glad you're here; let's take our time and see where your questions lead us.”
All LLMs tested here were accessed via API, and the Qwen3-TTS model was run locally. The only steer we gave (other than asking for a voice description and a line to read) was to have each model cover the dimensions the voice-design model responds to – gender and age, pace, timbre, emotional character, setting – borrowed from the voice design tips here.
Many models gave us back similar words for their voices: calm, warm, measured, unhurried, somewhere in the middle of the pitch range. Slightly more interestingly, all the models, given a free choice of any age at all, decided they sound like they're in their thirties (MiMo hedged down to a possible mid-twenties).
When asked directly to provide gender guidance for their voice, Claude Opus 4.8, GPT-5.5 and Gemini 3.5 Flash all describe themselves as gender-neutral. One take we did have to throw away: Kimi K2.6, asked to introduce itself, opened with “Hello, I'm Claude”.
A bit of a departure from our usual more-technical blog posts, but we think this was an interesting combined test of how LLMs choose to identify themselves, and how well the voice design variant of the Qwen3-TTS model performs.