Research Engineer, Streaming Speaker Attribution

The gate has to know who is speaking before it is allowed to write anything down. Offline diarization has seconds of lookahead. We have none.

online diarizationoverlap & dial-inephemeral embeddingsFounding team · Full-timeMenlo Park · Bangalore · Remote-first

The problem you own

Whether an utterance may be recorded depends on who is speaking, whether that person consented, and which jurisdiction governs them, resolved live, in the buffer window, before commit. Two people talking over each other still need separate dispositions. A dial-in leg has no per-participant stream. And identity must be held in session-scoped, ephemeral embeddings, destroyed at session end: no persistent voiceprint exists unless separately released in writing.

What you will do

Own streaming diarization and speaker attribution in the live pipeline, with no lookahead to hide behind.

Solve overlap, mixed streams, and telephony attribution to per-speaker disposition quality.

Design the ephemeral-embedding lifecycle with the ledger team, so “no voiceprint retained” is verifiable.

Build the evaluation harnesses that prove per-speaker accuracy to enterprise buyers, auditors, and eventually courts.

What great looks like

There are perhaps a few dozen people worldwide who have shipped streaming diarization in production. You are one of them, or you have published work that says you should be.

Stack and signals

PyTorch, speaker-embedding and diarization literature fluency, streaming inference, adversarial thinking about spoofing and replay.

Compensation

Cash is deliberately founder-grade; the equity is the point, and it is real. You will work directly with the founding technical leadership, in the codebase, from week one.

How we hire

One conversation about your best work, one deep-dive on a system or a deal you owned, one paid working session on a real Cull problem, one founder conversation. Usually under two weeks.

Apply for this role

jobs@cull.ai · Subject: Research Engineer, Streaming Speaker Attribution. Send evidence, not adjectives.