LiveKit Logo

LiveKit

Senior Software Engineer - Reliability, Infrastructure, and Tooling

Reposted Yesterday
In-Office or Remote
Hiring Remotely in Canada
Senior level
In-Office or Remote
Hiring Remotely in Canada
Senior level
The Senior Infrastructure Engineer will manage and scale LiveKit's core infrastructure, focusing on reliability and performance, while working on Golang code and automation for distributed systems.
The summary above was generated by AI
About LiveKit

LiveKit is building the infrastructure layer for the agentic era of computing. Our platform gives developers everything they need to build, test, deploy, scale, and observe AI agents in production. Founded in 2021, LiveKit powers voice and agentic AI applications for OpenAI, Salesforce, Spotify, Meta, and tens of thousands of other developers, collectively facilitating billions of calls each year.

About This Role

We’re hiring Senior Software Engineers to join our team focused on infrastructure development and reliability engineering. This is not an “ops” team by any means. We partner with product dev teams to co-design and develop along with them to ensure that LiveKit systems are reliable, maintainable, and secure. We work on internal tooling to provide a smooth experience for our product dev teams to own and run their workloads on top of our infrastructure and to meet our strict reliability requirements.

Like all teams, we own the ops and maintenance for the systems we work on, but where possible we automate away what we can. Our team also facilitates a healthy oncall rotation and incident management practices, but the rotation is shared with product dev team members to ensure the important production perspective that oncall provides isn’t isolated to just our team.

We support the full range of LiveKit products which provide a fascinating landscape of problems to be solved because they are much more demanding than a simple web app. It includes real time media workloads, hosting of customer agent code in secure sandboxes, and advanced networking requirements, all of which keep us on our toes.

You'll Thrive Here If You:
  • You are tenacious with investigating tricky system level issues.

  • You like the art of observability including quantifying reliability and visualizing it efficiently.

  • You understand the delicate balance between moving fast now and moving fast later.

  • You are able to communicate effectively with partner teams and tactfully handle sometimes contentious topics.

  • You get satisfaction from clean, DRY, error-resistant configuration even when the underlying systems are complex and diverse.

  • You look at the world in terms of signals and control systems.

What You'll Do
  • Ramp on LiveKit's global architecture — CockroachDB, NATS, Nebula, Kubernetes — and map where reliability debt lives

  • Ship product SRE work directly in the product codebase: load balancing, load shedding, instrumentation, scalability, efficiency

  • Build and extend common tooling so product teams can self-service reliability without Infra as a bottleneck

  • Participate in the on-call rotation and help resolve recurring reliability patterns

  • Bring informed systems opinions that improve how the team makes architectural decisions

Who You Are
  • Experience building non-trivial applications (high concurrency, complex control loops, etc).

  • Strong experience with Kubernetes (or equivalent, Borg, etc).

  • Experience with Linux internals and networking.

  • Experience making use of observability tools to debug tricky problems.

  • Experience running large scale globally distributed systems and working with the complex configuration management problems and technical debt that come along with it.

  • Experience with complex production incident handling.

  • Experience running open source tooling (e.g. Kafka, Clickhouse, etc).

Nice to Have
  • Data engineering and analytics.

  • Global layer 3 networking.

  • Experience working with systems that handle long lived load like media.

  • Google SRE or equivalent high-scale background.

  • Dealing with compliance frameworks (e.g. PCI).

Our Commitment to You
  • An opportunity to build something truly impactful to the world

  • Contribute to open source alongside world-class engineers

  • Competitive salary and equity package

  • Health, dental, and vision benefits

  • Flexible vacation policy

LiveKit is an equal opportunity employer and does not discriminate on the basis of any characteristic protected by applicable law. If you require a reasonable accommodation during the application or interview process, please contact [email protected].

Similar Jobs

Yesterday
In-Office or Remote
Canada
Senior level
Senior level
Artificial Intelligence • Information Technology • Internet of Things
Design, implement, and manage LiveKit's infrastructure, focusing on performance, reliability, and automation while ensuring effective incident management and coordination with vendors.
Top Skills: Container OrchestrationGoKubernetesLinux
Yesterday
Remote or Hybrid
Senior level
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Leads strategy, delivery, governance, adoption, and lifecycle management for agentic AI products across Pfizer’s regulated manufacturing network. Defines use cases, roadmaps, evaluation frameworks, guardrails, success metrics, and value-realization models while partnering with manufacturing, engineering, data, quality, cybersecurity, and business teams. Oversees pilots, production deployment, monitoring, continuous improvement, responsible AI, and scaling of intelligent manufacturing capabilities.
Top Skills: Agentic AiAICloud Data PlatformsHistoriansLimsLlmopsMesMlopsPlant ConnectivityScadaSerialization
Yesterday
Remote or Hybrid
Entry level
Entry level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Build, prototype, evaluate, deploy, and improve agentic AI products for Pfizer’s regulated manufacturing network. Own agent backlogs, user discovery, experimentation, adoption, success metrics, governance, responsible AI, monitoring, and lifecycle operations. Partner with engineers, data scientists, manufacturing SMEs, Quality, Cybersecurity, and business stakeholders to convert manufacturing knowledge and user needs into production-grade intelligent agents that improve productivity, decision-making, and operational performance.
Top Skills: Advanced AnalyticsAgentic AiAgentic Ai OrchestrationAgile/ScrumAi Agent FrameworksAi Evaluation FrameworksAi GuardrailsCloud Data PlatformsCsvHistoriansHuman-In-The-Loop OversightLarge Language Models (Llms)LimsLlmopsMesMlopsOt/AutomationPrompt EngineeringRetrieval-Augmented Generation (Rag)Scada

What you need to know about the Toronto Tech Scene

Although home to some of the biggest names in tech, including Google, Microsoft and Amazon, Toronto has established itself as one of the largest startup ecosystems in the world. And with over 2,000 startups — more than 30 percent of the country's total startups — Toronto continues to attract new businesses. Be it helping entrepreneurs manage their finances, simplifying business operations by automating payroll or assisting pharmaceutical companies in launching new drugs, the city's tech scene is just getting started.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account