Engineer, Real-Time Inference & The Gate
Classify every moment against compiled policy inside the buffer horizon, and choose: release to storage, or irrecoverably discard. Before the write, always.
The problem you own
The gate sits in front of the write. Audio lands in a transient in-memory buffer, attribution resolves, and your engine must produce a capture determination per speaker, keep, discard, redact, withhold, alert, route, before the buffer horizon expires. A fault must resolve toward less retention, never more: fail-closed is a property of your design, not a promise in a policy page.
What you will do
Own the model-serving stack for live classification: batching, small-model routing, distillation, structured outputs.
Enforce the latency budget as an engineering contract, with regression gates that block merges.
Design the fail-closed semantics: what suspends, what continues, and how degraded intervals are journaled and provable.
Manage GPU scheduling and unit economics so accuracy never quietly buys itself with cost.
What great looks like
You treat p99 latency as a design input, and you have owned production model systems where a wrong or slow answer had consequences someone escalated.
Stack and signals
Python, an inference stack you can defend, CUDA-adjacent comfort, evals as code.
Compensation
Cash is deliberately founder-grade; the equity is the point, and it is real. You will work directly with the founding technical leadership, in the codebase, from week one.
How we hire
One conversation about your best work, one deep-dive on a system or a deal you owned, one paid working session on a real Cull problem, one founder conversation. Usually under two weeks.
jobs@cull.ai · Subject: Engineer, Real-Time Inference & The Gate. Send evidence, not adjectives.