The stack, in one sentence
Microphone or file in. A speech model turns it into text. A language model summarises, classifies, or decides what to say next. A TTS model speaks, or you skip that step and keep the notes.
That pipeline is the easy part. The wrapper is everything around it: recording in the browser, job status, stored audio, timestamps, credits so a long call cannot bankrupt you, and a next action — copy the summary, book the meeting, hand off to a human.
If you only expose a “make it talk” box, you have a demo. If you own the job — voiceover, meeting notes, inbound qualification — you have a product.