AI voices can now laugh, sigh, and cry on command like… · M&A Beginners 🎓
| View this email in your browser |
![]() Models & Agents for BeginnersAI explained simply — for beginners and teens.
|
🎧 Today's episode Episode 117 · AI voices can now laugh, sigh, and cry on command like real actors. 2026-07-29 ▶ Listen now |
The Big StoryFish Audio released S2.1 Pro, a new voice AI that responds to plain-text directions like [nervous laugh] or [long sigh] to create emotional performances instead of flat readings. The same script run through Fish Audio, ElevenLabs, and Cartesia produced noticeably different results, with Fish Audio showing more natural pauses, breathing, and intonation. Think of it like giving stage directions to an actor in a play — you type the emotion in brackets and the AI performs it, adding pauses, breathing, and tone that match what a real person would do. The model also lets you control emphasis on individual words and works in over 80 languages while streaming audio fast enough for live conversations. It reaches the first audio in roughly 90 milliseconds, which keeps rhythm intact even when the conversation shifts quickly or gets interrupted. This shift matters because voice AI used to focus only on sounding realistic, but now the goal is making it expressive enough for stories, characters, or helpful assistants that feel natural during back-and-forth chats. For students, this could mean turning a history report into a dramatic reading with different voices and emotions, or creating podcast-style projects without recording equipment. For anyone who makes videos or games, it opens the door to characters that react with real feeling instead of robotic delivery. The practical side is that you can clone a voice from just 15 seconds of audio and run it at one-sixth the cost of ElevenLabs. HeyGen has already integrated the model, and the open-weights versions can be self-hosted. You might find this changes how you interact with AI on your phone in the future, since expressive voices could make homework helpers or creative tools feel more like talking to a friend. Right now you can try it yourself at the Fish Audio site linked in the original post — start with the free tier, type a short script, and add tags like [excited] or [whisper] to hear the difference. Source: x.com Explain Like I'm 14You know how when you're texting a friend and your phone suggests the next word based on what you've already typed? Now picture that same idea but for sound instead of text. The AI has learned patterns from thousands of real voice recordings, so when you add a tag like [nervous laugh], it doesn't just guess the next sound — it pulls from the patterns that match nervous laughter in actual human speech. Next, the model breaks your whole sentence into tiny pieces and decides exactly where to place the laugh, how long it should last, and how the words around it should rise or fall in pitch. It also keeps track of the rhythm so the voice still flows naturally even when you interrupt or change topics mid-sentence. The result is that the AI isn't just reading words — it's performing them the way a voice actor would after getting the same directions. This is why the output feels alive instead of mechanical. Cool Stuff & Try ThisMake a digital twin of yourself from a 10-second video Avatar X creates a moving version of you that captures the way your eyes move, how you shrug, and even micro-expressions that appear when you talk. Most older avatar tools only copy your face shape and lip movements, but this one learns your full identity from a short clip so the result feels like watching yourself on a video call. It preserves the way you naturally move and express yourself rather than forcing every sound into lip-sync. It's exciting because it removes the need for long recording sessions or expensive equipment — anyone could soon make a version of themselves for school presentations or creative projects. The model handles non-verbal moments like laughing, crying, yawning, or sighing without breaking character, and expressions spread across the whole face and body instead of staying limited to the mouth. Quality stays consistent from the first second to the last, unlike earlier tools that degrade over longer clips. Go to the link in the original post, upload 10 seconds of yourself speaking naturally, and watch how the eyes and expressions stay consistent even during longer clips. Watch AI treat a cheese chart like serious data Ethan Mollick shared an infographic about cheese that uses a carefully tested color scale so it stays readable in both light and dark modes. The chart uses a six-step amber ramp that passed every contrast check in a dataviz validator. The panel stays cave-dark in both themes because white cheese has no contrast against a pale background, so the light-mode version failed and the dark inset became the honest solution. The fun part is seeing how an AI model dives deep into the details of the chart even though the topic is lighthearted. This shows how AI can take any visual seriously and help spot design improvements. Head to the link in the post to see the chart and notice how the dark background keeps the white cheese areas from blending in. Source: x.com Quick BitsAnthropic backs a petition for slower AI progress The company and its leaders signed a petition asking the field to create tools that deliberately slow down the fastest AI systems so society has time to prepare. It connects to their recent research showing AI could improve itself in loops, which raises questions about pacing. The petition was signed by the CEO, several co-founders, and senior staff. Source: x.com OpenAI shares a free security scanner for code The tool checks repositories for problems, tracks fixes over time, and can be added to automated workflows. It's an early release meant for anyone who wants to keep projects safer as AI tools grow more common. You can install it with the command npm install @OpenAI/codex-security or run it directly using npx @OpenAI/codex-security@latest --help, and the full source plus documentation lives on GitHub at the link in the post. |
💬 Reply to this email — Patrick reads every one. Share: X · LinkedIn · WhatsApp Forwarded this email? Subscribe here — it's free. |
📺 Watch on YouTube · 📝 Read the blog · 🖼 Free image gallery (CC BY-SA) · 📊 Data Hub & Story Trackers · 🧭 Start Here Nerra Network · AI-narrated voice (Grok TTS) · Editorial by Patrick You're receiving this because you subscribed to Models & Agents for Beginners on nerranetwork.com. |
| Issue #117 · Models & Agents for Beginners · Jul 29, 2026 |
