The word "Gemini" in bold dark letters on a pink and dark blue abstract background with star shapes.

Google adds Gemini 3.5 audio tools that clean up transcripts and handle jargon

Google’s Gemini audio update adds cleaner transcription, jargon support, and 85+ languages while Gemini 3.5 Pro remains delayed.

In short

Google has released new Gemini Audio models, led by Gemini 3.5 Transcribe, which can remove filler words, handle specialist jargon, and support more than 85 languages. The rollout improves voice features across macOS, Android dictation, and the Gemini API while users still await Gemini 3.5 Pro.

  • Gemini 3.5 Transcribe removes filler words, supports custom vocabulary, and adds word-level timestamps.
  • Google says the new transcription model improves multilingual accuracy and reduces wording errors compared with Chirp 3.
  • Gemini 3.5 Live improves real-time voice handling, including interruptions, language recognition, and visual processing.
  • The update is rolling out now on macOS, select Android dictation markets, and in public preview for developers.
  • The release arrives while Google’s promised Gemini 3.5 Pro model is still overdue.

Google has begun rolling out new Gemini Audio models, including Gemini 3.5 Transcribe, a transcription system that can remove filler words, adapt to specialist terminology, and work across more than 85 languages. The update matters because it gives Gemini’s voice features more accurate, more polished transcripts at a time when Google is still overdue to ship its Gemini 3.5 Pro model.

The company says the new audio stack is designed to improve voice-controlled AI when conversations are noisy, interrupted, or multilingual, while also making the output easier to edit and use in real workflows. For developers and app users, the rollout is another sign that Google is pushing Gemini further into practical productivity tools, not just conversational assistants.

What Google launched in Gemini Audio

Google’s latest update centers on three models: Gemini 3.5 Live, Gemini 3.5 Live Experimental, and Gemini 3.5 Transcribe. Together, they expand Gemini’s audio capabilities with better transcription, stronger speech recognition, and more resilient live voice handling.

The headline addition is Gemini 3.5 Transcribe, which Google describes as a major upgrade over its earlier transcription engine, Chirp 3. The company says the new model performs better on multilingual speech and reduces wording errors, especially in situations where previous systems would stumble.

That includes audio with overlapping speech, background noise, or mid-sentence interruptions. Google is positioning the release as a practical improvement for dictation, note-taking, customer workflows, and any app that depends on clean voice input.

Why the timing stands out

The launch arrives while users are still waiting for Gemini 3.5 Pro, the flagship model Google had previously said would arrive in June. That delay has made the audio update notable in its own right, because it shows Google continuing to ship parts of the Gemini family even as the broader model roadmap remains unfinished.

In other words, Google is using audio to keep momentum going. Rather than waiting for the full release of its next major general-purpose model, the company is delivering specialized tools that can be folded into apps and developer products now.

How Gemini 3.5 Transcribe works

Gemini 3.5 Transcribe is built to turn spoken language into text with more context awareness than a basic speech-to-text system. One of its most useful features is the ability to strip out conversational filler such as “um” and “uh,” which can make a transcript feel cleaner and more readable without requiring manual editing.

The model can also format text automatically, which should help users turn raw speech into notes, drafts, or records that are closer to finished copy. Google says users can even “edit naturally with just your voice,” suggesting the system is intended for interactive dictation rather than passive capture.

Another important capability is custom vocabulary support. Users can feed the model specialized terms, brand names, technical jargon, or unusual spellings so the transcription engine can learn what to preserve. That should be especially valuable for fields where standard speech models often fail, including healthcare, software, law, manufacturing, and research.

Who benefits most from the custom vocabulary feature?

Anyone who regularly uses specialist language stands to benefit, especially professionals who cannot afford to have important terms edited incorrectly. Google says the model can adapt to unique spelling requirements and jargon, which reduces the need for manual correction later.

For example, a developer dictating product names, a clinician using medical terminology, or a journalist transcribing proper nouns may all get cleaner results with less post-processing.

What is different about Gemini 3.5 Live?

Gemini 3.5 Live is Google’s real-time speech model for voice interactions, and the company says it is better at handling interruptions, identifying languages, and processing visual inputs while a conversation is underway. Those changes matter because live AI assistants rarely operate in ideal conditions.

People interrupt themselves. Background noise cuts in. Speakers switch languages. Sometimes the assistant needs to understand what it is seeing as well as what it is hearing. Google says 3.5 Live is meant to perform more reliably in those messy, real-world scenarios.

The more experimental version, Gemini 3.5 Live Experimental, goes further by narrating its own reasoning step by step in real time as it works through more difficult tasks. That makes it less of a pure voice dictation engine and more of a visible problem-solving assistant, though Google is clearly treating that capability as early-stage.

How does the experimental live model differ?

The experimental build is aimed at complex reasoning and more transparent execution. Instead of silently producing an answer, it explains its progress as it goes, which may help users follow what the model is doing and catch mistakes earlier.

That kind of feature can be useful for debugging, task planning, or multi-step assistance, but it also signals that Google is testing a more agent-like style of interaction. The tradeoff is that experimental features often change quickly and may not be as reliable as the stable release.

Why the transcription upgrade matters for everyday users

Transcription tools are often judged less by benchmarks than by annoyance. If a product can clean up filler words, recognize domain-specific terms, and avoid mangling speech in noisy environments, it saves time and reduces friction in ordinary workflows.

That is what makes Gemini 3.5 Transcribe interesting. It is not just about converting audio to text; it is about turning spoken content into something immediately useful. A transcript that does not require extensive manual cleanup is far more valuable for meeting notes, interviews, drafts, voice memos, and accessibility use cases.

Google is also emphasizing multilingual performance. Support for more than 85 languages puts the company in a stronger position to compete in international markets and to support mixed-language conversations, which are increasingly common in global workplaces and consumer apps.

Where the new Gemini Audio models are available

Google says the rollout starts immediately for English-language users on the Gemini app for macOS. The company is also bringing the updates to its Rambler dictation feature on Android, though that rollout is limited to select countries and languages.

For developers, the models are available in public preview through the Gemini API via AI Studio and Antigravity. That means app makers can begin testing the new transcription and live audio capabilities before they are fully hardened for broader deployment.

Chrome support is also on the way, according to Google, which could make the audio stack more broadly useful inside the company’s browser ecosystem. That would extend Gemini’s voice features into another major surface where users already spend a large share of their online time.

Model Main purpose Key capabilities Availability
Gemini 3.5 Transcribe Speech-to-text and editing Filler-word removal, custom vocabulary, multilingual transcription, speaker attribution, word timestamps Public preview via Gemini API; rolling out now in supported apps
Gemini 3.5 Live Real-time voice interaction Better interruption handling, language recognition, live visual processing Rolling out now
Gemini 3.5 Live Experimental Advanced live reasoning Step-by-step real-time narration of reasoning while solving complex tasks Rolling out now in experimental form

What Google says Transcribe can do

Google says Gemini 3.5 Transcribe can attribute speech among up to three speakers in pre-recorded audio and provide word-level timestamps, which should make it more useful for interview clips, meetings, and other structured recordings. Those features also help users find and cite specific moments quickly.

The model’s ability to preserve speaker separation is especially important for teams and creators who depend on accurate transcription for editing, publishing, or documentation. A transcript that clearly shows who said what is easier to search, verify, and reuse.

Google says the model is designed to let users edit naturally by voice and to cut down on the manual work of cleaning transcripts, while also preserving specialized language that might otherwise be lost.

That framing suggests the company sees transcription not as a back-end utility but as an interactive layer in its AI product strategy. The goal is to make speaking to the computer feel more like editing with the computer.

How does it compare with earlier Google transcription tech?

Google says Gemini 3.5 Transcribe is a significant step up from Chirp 3, the earlier transcription model. The improvements it highlights are multilingual accuracy and fewer wording mistakes, two metrics that usually separate a merely usable tool from one that feels trustworthy.

That comparison matters because transcription systems are often only as good as their weakest cases: names, accents, technical vocabulary, and background chatter. If Gemini 3.5 can reduce errors in those areas, the upgrade could be meaningful even without flashy consumer-facing changes.

The new model also extends the Gemini family in a more specialized direction. While flagship chat models usually grab attention, transcription can be far more operationally useful for businesses and creators who need consistent output at scale.

Why Google is emphasizing voice now

Voice remains one of the clearest ways to make AI feel immediate and useful. Speaking is faster than typing for many tasks, and AI can convert that speech into summaries, drafts, or actions. Google’s update suggests it is still betting that voice will become one of the main interfaces for everyday AI use.

That strategy also makes sense for Google’s product ecosystem. The company has search, Android, Chrome, productivity software, and a growing developer platform; improved voice capabilities can be inserted across all of them.

There is also a competitive angle. As rivals push their own assistant and transcription tools, Google needs Gemini to feel useful beyond generic chat. Cleaner transcripts, better multilingual support, and more responsive live voice features all help move the product in that direction.

What this means for developers

Developers are getting early access through the Gemini API, which is often the most important part of any launch like this. Public preview lets app builders test whether the new models improve their own products before committing to a full integration.

That could matter for meeting platforms, dictation tools, note-taking apps, customer-service systems, and accessibility software. If Gemini 3.5 Transcribe proves more accurate in noisy or multilingual conditions, it may become a strong foundation layer for third-party tools.

  • Meeting apps could generate cleaner notes and speaker labels.
  • Productivity tools could use voice to revise text faster.
  • Developer platforms could expose custom jargon dictionaries to users.
  • Mobile apps could improve real-time dictation for multilingual users.

What remains unresolved

Despite the new audio models, Google still has to answer broader questions about Gemini’s release cadence. The company’s overdue Gemini 3.5 Pro launch will likely remain the main benchmark for how quickly Google can deliver its flagship AI ambitions.

Until that model arrives, smaller releases like the audio update will be read in two ways: as useful product progress, and as evidence that the overall roadmap is moving more slowly than promised.

There is also the practical matter of reliability. Transcription systems can look strong in controlled demos but struggle in daily use, especially when users speak casually, switch languages, or work in loud environments. The coming weeks will determine whether Google’s accuracy claims translate into real-world trust.

Timeline of the Gemini Audio rollout

The sequence below shows how Google’s audio announcements fit into the broader Gemini release picture.

Date / period Event Why it matters
June 2026 Google said Gemini 3.5 Pro would launch The flagship model has yet to arrive, raising pressure on Google’s roadmap
August 26, 2026 Gemini Audio update announced Google ships specialized audio models instead of waiting for the main release
Starting now Rollout begins in macOS Gemini app, Android Rambler dictation, and API preview Users and developers can begin testing the new transcription and voice tools immediately
Soon Chrome support promised Could widen Gemini’s reach inside Google’s browser ecosystem

The bigger picture for Gemini

Google’s audio release reinforces a broader pattern in its AI strategy: ship useful building blocks even when the marquee model is not ready. That approach helps keep Gemini visible and gives developers something tangible to work with.

It also shows how quickly AI products are fragmenting into specialized capabilities. Rather than relying on a single large model to do everything, companies are adding focused systems for transcription, live conversation, reasoning, and multimodal input.

For users, that can be a good thing if the features are reliable and integrated cleanly. It means better voice dictation, cleaner transcripts, and more adaptive assistants. For Google, it is a chance to make Gemini feel indispensable across devices and workflows.

But the company still faces the same challenge it has had throughout the Gemini rollout: proving that its AI products are not only ambitious, but also consistently available, practical, and better than what came before.

Key details at a glance

Here are the main facts about the update:

  • New model: Gemini 3.5 Transcribe
  • Other models: Gemini 3.5 Live and Gemini 3.5 Live Experimental
  • Language support: More than 85 languages
  • Notable features: Filler-word removal, custom vocabulary, speaker attribution, timestamps
  • Availability: Rolling out now for macOS Gemini users, select Android dictation markets, and developers via public preview

For now, the update gives Gemini a more practical voice toolkit and a better transcription engine. Whether that is enough to quiet the focus on the missing Gemini 3.5 Pro will depend on how quickly Google can turn these specialized gains into broader platform momentum.

Frequently asked questions

What is Gemini 3.5 Transcribe?

Gemini 3.5 Transcribe is Google’s new speech-to-text model for Gemini Audio. It is designed to produce cleaner transcripts, remove filler words, support custom jargon, attribute speakers in pre-recorded audio, and provide word-level timestamps.

When is Gemini Audio rolling out?

Gemini Audio is rolling out now. Google says the update is available in English for all macOS Gemini app users, on Android through Rambler dictation in select countries and languages, and in public preview for developers through the Gemini API.

How is Gemini 3.5 Live different from Transcribe?

Gemini 3.5 Live is built for real-time voice interaction, while Transcribe focuses on converting recorded speech into polished text. Live is better at interruptions, language recognition, and visual processing, whereas Transcribe emphasizes cleaner dictation and editing.

Why does this update matter for users?

This update matters because it makes voice input more useful in everyday work. Cleaner transcripts, fewer filler words, better jargon handling, and stronger multilingual support can save time for meetings, interviews, dictation, and note-taking.

Is Gemini 3.5 Pro available yet?

No, Gemini 3.5 Pro was promised for June but has not been released yet. Google’s new audio models arrived while users are still waiting for that flagship launch, making the audio update more significant in the meantime.

Share this 🚀