● AN INDEPENDENT LAB FOR FULL-DUPLEX AUDIO INTELLIGENCE
Full-duplex models can already hear, think and speak at once — in English. We extend the best open ones: our first mission is a multilingual PersonaPlex, built on NVIDIA's open weights.
THESIS
Voice AI today is a walkie-talkie. It waits for you to finish, transcribes, thinks, then performs a reply.
Human conversation doesn't take turns. It overlaps — interruptions, backchannels, silences that mean something.
Full-duplex models finally live inside that timing. We are teaching them to do it in every language.
The duplex breakthrough already happened — PersonaPlex, Moshi and their kin listen and speak in one continuous stream. Rebuilding that would burn compute on a solved problem. We start from open weights and spend every GPU-hour on what's missing.
Timing is cultural. How long a polite pause lasts, when overlap is warmth and when it's rudeness, what a backchannel sounds like — all of it differs by language. A multilingual duplex model has to learn each language's rhythm, not just its words.
Multilingual conversational speech, synthetic overlapping dialogue, and post-training that adds languages without breaking the duplex behaviours already there — with yield latency and backchannel timing benchmarked per language.
Half the economy runs on conversation — most of it not in English. Full-duplex makes machine voice employable; multilingual makes it global.
Agents that can be interrupted, corrected and talked over — and still hold the thread, in the customer's own language.
NEAR-TERMListening and speaking at once is the job description. Translation that keeps pace with the speaker — not after them.
NEAR-TERMReal-time conversational support for people who speak, hear or process differently. Timing is dignity: no dead air, no being talked over.
PILOTIntake and triage that listens the way a nurse does — probing, confirming, reassuring while the patient is still talking.
PILOTRobots, vehicles and machines that coordinate with people by voice, in places where hands and eyes are already busy.
RESEARCHBuilding toward first pilots with design and compute partners. If that could be you — hello@instituteofthinking.com
Real-time audio systems. Streaming inference at conversational latency.
OPENINGPost-training duplex speech models. Multilingual conversational data.
OPENINGNo forms, no cover letters. Send the most interesting thing you've built to hello@instituteofthinking.com