AN INDEPENDENT LAB FOR FULL-DUPLEX AUDIO INTELLIGENCE

Machines that listen while they speak.

Full-duplex models can already hear, think and speak at once — in English. We extend the best open ones: our first mission is a multilingual PersonaPlex, built on NVIDIA's open weights.

THE APPROACH ↓ HELLO@INSTITUTEOFTHINKING.COM
FIG. 1 — THE THREE-BODY PROBLEM t = 000.0 · E = −1.29 ● LIVE
m₁ — a voice m₂ — a voice m₃ — the model
CLICK TO PERTURB
FIG. 1 — A conversation is a three-body problem: two voices and a model in mutual orbit, plotted pixel by pixel. Simple rules, never repeating. Click near a body to perturb it.

THESIS

Voice AI today is a walkie-talkie. It waits for you to finish, transcribes, thinks, then performs a reply.

Human conversation doesn't take turns. It overlaps — interruptions, backchannels, silences that mean something.

Full-duplex models finally live inside that timing. We are teaching them to do it in every language.

01 — APPROACHENHANCE, DON'T REBUILD
A.1

Stand on open models

The duplex breakthrough already happened — PersonaPlex, Moshi and their kin listen and speak in one continuous stream. Rebuilding that would burn compute on a solved problem. We start from open weights and spend every GPU-hour on what's missing.

A.2

Multilingual is not translation

Timing is cultural. How long a polite pause lasts, when overlap is warmth and when it's rudeness, what a backchannel sounds like — all of it differs by language. A multilingual duplex model has to learn each language's rhythm, not just its words.

A.3

Post-training as a craft

Multilingual conversational speech, synthetic overlapping dialogue, and post-training that adds languages without breaking the duplex behaviours already there — with yield latency and backchannel timing benchmarked per language.

FIG. 2 — A DUPLEX EXCHANGE, SIMULATED LAST YIELD · 180 MS
HUMAN MODEL
FIG. 2 — Both channels are open the whole time. The model backchannels while listening and yields in ~180 ms when you barge in. Press INTERRUPT to try it.
02 — APPLICATIONSWHERE DUPLEX BECOMES AN ECONOMY

Half the economy runs on conversation — most of it not in English. Full-duplex makes machine voice employable; multilingual makes it global.

B.1

Customer operations

Agents that can be interrupted, corrected and talked over — and still hold the thread, in the customer's own language.

NEAR-TERM
B.2

Simultaneous interpretation

Listening and speaking at once is the job description. Translation that keeps pace with the speaker — not after them.

NEAR-TERM
B.3

Accessibility

Real-time conversational support for people who speak, hear or process differently. Timing is dignity: no dead air, no being talked over.

PILOT
B.4

Healthcare front line

Intake and triage that listens the way a nurse does — probing, confirming, reassuring while the patient is still talking.

PILOT
B.5

Embodied systems

Robots, vehicles and machines that coordinate with people by voice, in places where hands and eyes are already busy.

RESEARCH

Building toward first pilots with design and compute partners. If that could be you — hello@instituteofthinking.com

03 — NOTESWORKING NOTES, PUBLISHED IRREGULARLY
№ 001

A multilingual PersonaPlex — the mission

SEP 2026 — FIRST NOTE
№ 002

Yield latency across languages

IN DRAFT
№ 003

The three-body problem of conversation

IN DRAFT
04 — CAREERSFOUNDING ROLES
C.1

Founding research engineer

Real-time audio systems. Streaming inference at conversational latency.

OPENING
C.2

Founding researcher

Post-training duplex speech models. Multilingual conversational data.

OPENING

No forms, no cover letters. Send the most interesting thing you've built to hello@instituteofthinking.com