Senior Software Engineer: Applied AI (Voice Agents & ML Systems) The pitch
We build and operate production AI voice agents that hold real phone conversations in a regulated healthcare setting, plus the machine learning and LLM pipelines around them. This is one seat that spans four disciplines that rarely come together: realโtime systems, LLM engineering, traditional machine learning, and serious cloud infrastructure, all in production, all with real consequences. If you are the kind of engineer who gets restless doing one thing, this role is the opposite problem.
What youโll work across
- Streaming, lowโlatency speechโtoโspeech systems built on modern LLMs
- Telephony and realโtime media (call control, live audio streaming)
- Audio handling and the quirks of real human conversation (interruptions, timing, noise)
- Concurrency on a latencyโsensitive path, where p99 matters and a stall is something a caller hears
- Wrapping nondeterministic models in deterministic control so they behave reliably in production
- Multiโmodel pipelines, prompt design, and cost/latency budgeting
- Evaluation harnesses, including LLMโasโjudge and automated agentโtestsโagent approaches
- Agentic tooling that gives AI systems safe, structured access to infrastructure
Traditional (nonโLLM) machine learning
- Endโtoโend ML pipelines: feature engineering, model training, and scheduled inference
- Imbalanced, messy realโworld data; calibration and explainability for nonโtechnical consumers
- Turning research notebooks into reproducible, auditable production pipelines
Cloud and infrastructure
- Infrastructure as code across multiple environments (we run on AWS)
- Managed compute, data, streaming, and orchestration services
- Security engineering in a regulated setting: encryption, leastโprivilege access, strict dataโhandling discipline
- Observability and telemetryโdriven debugging, tracing a production issue from a metric anomaly to root cause
Plus
Occasional fullโstack work on internal tools, and an engineering workflow that leans heavily on AI coding assistants, with human accountability for every change.
What youโll actually do
- Ship and debug code on a live, realโtime voice pipeline where latency and correctness are userโfacing
- Design control systems around LLMs: guardrails, budgets, watchdogs, safe fallbacks
- Build and operate LLM evaluation and batchโanalysis pipelines
- Own traditional ML workflows from data to scheduled production inference
- Trace production issues from a metric anomaly to root cause, including building the evidence when the cause is a vendor
Mustโhaves
- 7+ years building and operating production backend systems, with strong generalโpurpose programming skills (we work primarily in Python)
- Experience running distributed systems in the cloud; comfortable debugging from telemetry to root cause
- Handsโon production experience with LLMs or generative AI (any provider or framework), plus the judgment to know when not to use a model
- Working fluency across the traditional machine learning lifecycle (you productionize; you do not need to publish)
- Disciplined in a regulated environment: small, reviewable changes and careful handling of sensitive data
Niceโtoโhaves
- Realโtime media or telephony experience
- Frontโend / fullโstack ability
- ML pipeline experience, vector search, or embeddings
- Fluency with AI coding assistants (our workflows assume them, with human accountability for every change)
How we work
Smallest correct change wins. Every behavior change is validated against the live system. Evidence over opinion in debugging. Code review is rigorous. Safety and privacy gate everything.
This role is open only to US citizens and lawful permanent residents (Green Card holders). We cannot consider candidates who require visa sponsorship now or in the future, and we are unable to make exceptions of any kind.
How to apply
- Your LinkedIn profile URL
- A phone number where we can reach you
A resume is welcome but optional; the two items above are required.
#J-18808-Ljbffr