Gemini just dropped a dictation tool that turns messy… · M&A Beginners 🎓
| View this email in your browser |
![]() Models & Agents for BeginnersAI explained simply — for beginners and teens.
|
🎧 Today's episode Episode 147 · Gemini just dropped a dictation tool that turns messy voice notes into polished text while understanding what’s on your screen. 2026-08-27 ▶ Listen now |
The Big StoryImagine you’re voice-typing a school essay or a message to a friend, but you keep saying “um” and “like” and the words come out jumbled. Google’s new Gemini 3.5 Transcribe listens to your voice and turns it into clean, well-formatted text that actually makes sense. The model automatically filters out filler words, formats unstructured speech, and pairs with your screen context to execute voice commands. It supports 85-plus languages and can hand off tasks to other Gemini models through function calling. Think of it like having a super-patient friend who not only writes down what you say but also removes the awkward pauses, fixes grammar on the fly, and even looks at whatever app or document is open on your screen to understand the full picture. It works across more than 85 languages and can follow voice commands that relate to what you’re already looking at. In streaming mode the system achieves a 4.0 percent word error rate while cutting latency by 70 percent compared with the earlier Chirp 3 model. This matters because lots of people — students writing papers, creators drafting scripts, or anyone who thinks faster than they type — now have a tool that feels more like talking than typing. It lowers the barrier for turning ideas into finished work without getting stuck on spelling or structure. The same capabilities appear in the Gemini API, letting developers build custom dictation features inside their own apps via Google AI Studio or the Gemini Enterprise Agent Platform. For you personally, it could mean finishing homework voice notes in half the time or turning a rambling voice memo into a polished email. The model is already available to try in the Gemini app on macOS and through Gboard on Android devices. It is also coming soon to Google Chrome and Gemini Enterprise for Customer Experience. Just open the Gemini app, start a new chat, and tap the microphone to dictate something like a paragraph for an assignment — watch how it cleans up your speech and formats it neatly. You can also try it in any text field on Android by switching to Gboard and holding the microphone key. If you want to experiment with the API side, head to Google AI Studio and test a simple voice-to-text prompt without writing any code. Source: the-decoder.com Explain Like I'm 14You know how when you’re texting, your phone sometimes guesses the next word based on the sentence so far? Now picture that same idea, but instead of just one word it guesses an entire cleaned-up paragraph while also peeking at the app or document open on your screen. First it listens to the raw audio and breaks the sound into tiny chunks it can recognize as possible words. Then it uses the surrounding conversation or on-screen content as extra clues — like noticing you’re in an email app so it should sound more formal. Next it removes filler sounds such as “um” and “uh” because it has learned from millions of examples what polished writing usually looks like. Finally it rearranges the words into proper sentences and even suggests actions if you say something like “send this to Mom.” The system reaches a 4.0 percent word error rate in streaming mode because the extra screen context helps it pick the right words even when the audio is unclear. It also runs 70 percent faster than the previous Chirp 3 model by handing off complex tasks to other Gemini models through function calling. The result feels like the AI understood not just the sounds but the whole situation you’re in. That’s why the output comes out ready to use instead of needing tons of fixes afterward. Cool Stuff & Try ThisEarphones that let you chat with AI agents on the go Plaud’s new earbuds come with a special case that has its own internet connection built in, so you can talk to AI helpers without needing your phone nearby. They’re designed for quick voice conversations with AI agents while you’re walking, studying, or commuting. The buds are priced at $249 and use an eSIM-enabled case to stay connected wherever you go. Anyone who likes voice assistants but wants something more portable and always-connected should give them a look. The buds cost $249 and are available now. Once you have them, try asking the AI in the case to summarize a podcast you just heard or help brainstorm ideas for a project — all through voice only. You may need a parent’s help to set up the eSIM if you’re under 18. Source: techcrunch.com Quick BitsAnthropic shares real Claude chat data with outside researchers For the first time, Anthropic is letting university teams study how people actually use Claude in everyday conversations while keeping the chats private. Three research groups — Stanford’s Social and Language Technologies lab, Oxford’s Human Information Processing Lab, and METR — analyzed 250,000 aggregated Claude conversations from April and May 2026. Stanford’s SALT Lab found that over half of those conversations involved consequential tasks that affect other people or are hard to undo. This helps everyone understand the real impact of AI instead of just guessing. Source: anthropic.com Hugging Face launches an adorable tiny rollerskating robot duck The Microduck is a one-eyed biped robot under 10 inches tall that can roll around on tiny skates. It comes in cream, graphite, lavender, and sky blue and is available to preorder now for $399. Pollen Robotics, the Hugging Face company behind it, plans to start shipping before Christmas 2026. The little robot gives robotics fans a fun, small-scale project to play with without needing any programming experience. Source: theverge.com |
💬 Reply to this email — Patrick reads every one. Share: X · LinkedIn · WhatsApp Forwarded this email? Subscribe here — it's free. |
📺 Watch on YouTube · 📝 Read the blog · 🖼 Free image gallery (CC BY-SA) · 📊 Data Hub & Story Trackers · 🧭 Start Here Nerra Network · AI-narrated voice (Grok TTS) · Editorial by Patrick You're receiving this because you subscribed to Models & Agents for Beginners on nerranetwork.com. |
| Issue #147 · Models & Agents for Beginners · Aug 27, 2026 |
