Google's Gemini 2.0 Flash Live API now supports real-time, low-latency voice conversation, and developer Peter Yang used it to build Tabi, a working AI Japanese language tutor, in 6 discrete steps using the Gemini API through Google AI Studio.
The build is worth reading in full because the steps reveal the actual architecture decisions: how Live API handles streaming audio differently from standard text completions, how Yang structures system prompts for language coaching behavior, and how he integrates his open-source human-review skill from GitHub to keep the tutor's corrections accurate and grounded.
Voice as the primary human-computer interface is the thesis driving this project. Yang's argument is not abstract: he built the thing, shipped it, and documented the process in a way that makes replication straightforward. If you are building any real-time conversational AI product, the implementation specifics here are the closest thing to a practical blueprint currently available.
[WATCH ON YOUTUBE →]