Live Stream Date/Time: Thursday, October 1st, 2026 @ 9am Pacific.
Akshat Mandloi, co-founder of Smallest.ai, joins OpenCV Live to answer a question most of us have wondered on hold with customer service: why can you still tell it’s a bot?
His answer isn’t “the model’s too small.” After three generations of voice AI, less than 1% of the voice market is automated, and Akshat argues the problem is structural. Today’s agents listen, then think, then speak. People do all three at once, and interrupt each other while they’re at it.
He’ll walk through how the field got here, from the old ASR-to-LLM-to-TTS pipeline to full-duplex models that can hear while they talk, and why measuring “does this sound human” is still harder than it looks. Then the good part: how Smallest.ai built a speech model that scores 96% on Big Bench Audio and an agent that holds its own against frontier models at about a twentieth of the size, and where they’re heading next with a model that predicts a conversation instead of reacting to it.
Join our live stream and stick around for our giveaway of a free OpenCV University course to one lucky viewer.
- Watch on Patreon (DRM-Free, Ad-Free!)
- Watch on YouTube
- Watch on Apple Podcasts
- Watch on LinkedIn
- Watch on Zoom
- Watch on Twitch
Got a cool project of your own? Send it to us and you may be featured on a future episode.
The post How Machines Learned to Talk – OpenCV Live! 227 appeared first on OpenCV.