MaintainX Logo

MaintainX

Site Reliability Engineer

Posted One Month Ago
In-Office or Remote
Hiring Remotely in Toronto, ON, CAN
Mid level
In-Office or Remote
Hiring Remotely in Toronto, ON, CAN
Mid level
Partner with product and platform teams to improve reliability, observability, and developer autonomy. Design for reliability, establish standards, build shared tooling, mentor teams on SRE practices, and lead incident response and service health initiatives to scale platform resilience.
The summary above was generated by AI

MaintainX is a leading mobile-first work execution platform for industrial and frontline teams. More than 13,000 customers, including Duracell, McDonald's, Shell, DHL and Volvo, use MaintainX to cut unplanned downtime and run better operations, across 13.9 million managed assets and 79.5 million completed work orders.

In August 2026 MaintainX became part of Autodesk, joining Autodesk Operations Solutions, the organization unifying Autodesk's operations platform alongside Tandem, FlexSim and Fusion Operations. Autodesk's strategy is to converge design, make and operate into one continuous lifecycle: design an asset, build it, run it, then feed what you learn running it back into the next design. Autodesk had design and make. Operate is the phase that tells you what actually happened, and it is ours.

We’re looking for a Site Reliability Engineer to help advance MaintainX’s reliability, observability, and developer autonomy as we scale our platform.

In this role, you’ll partner closely with product and platform development teams to improve the stability, resilience, and operational readiness of our services. You’ll work alongside teams to design for reliability from the start, establish clear ownership and standards, and build shared tooling that enables teams to operate their services with confidence.

You’ll also contribute to company-wide initiatives that define how MaintainX approaches reliability software development, including observability standards, incident response practices, and service health metrics, helping the organization adopt proven industry practices at scale.

This role is well-suited for an developer who enjoys working across teams, influencing technical direction through strong development practices, and turning reliability principles into practical, scalable systems.

What You'll Do:

  • Assess service maturity and provide insights to development teams

  • Partner with development teams to implement observability best practices

  • Enable development teams to become autonomous with their service deployment, support, and infrastructure

  • Mentor developers on reliability practices, focusing on making them self-sufficient

  • Act as the bridge, ear and eyes of the Platform Division teams to drive tooling and practice adoption across development teams

About You:

  • Deep understanding of observability practices in a distributed system environment and how it influences system design and team behaviour

  • Practical experience with SRE concepts (SLOs, error budgets, incident management)

  • 3–5+ years in software development, SRE, DevOps, or production development roles with experience operating production systems

  • Proficient in cloud-native platforms and infrastructure-as-code concepts and tools

  • Working knowledge of at least one programming language (TypeScript/Node.js is a plus)

  • Excellent communication and collaboration abilities across technical and non-technical teams

  • Ability to translate complex reliability concepts into actionable guidance

  • You enjoy enabling teams to succeed independently and measuring success by reduced dependency on you

About Us:

MaintainX is committed to creating a diverse environment. All qualified applicants will receive consideration for employment without regard to race, colour, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, or veteran status.

 

Our mission is to keep the physical world running. Factories, fleets, hospitals and campuses stay up because the people who maintain them have tools worth using. That is what we build.

Compensation and benefits. Base pay is one part of the package. Depending on the role, compensation may also include commission, an annual bonus and equity. Benefits differ by country. For roles in the United States, Autodesk’s benefits are described at benefits.autodesk.com. For roles in Canada and other countries, the plan differs on health coverage, retirement and leave, and your recruiter will walk you through it.

Belonging. We take pride in a culture where everyone can thrive. More at autodesk.com/company/global-belonging. More on where this is going: Autodesk CEO Andrew Anagnost on building the future of connected operations, and AOS SVP Stephen Hooper on welcoming MaintainX to Autodesk.

Similar Jobs

8 Days Ago
Easy Apply
Remote
Canada
Easy Apply
Expert/Leader
Expert/Leader
Cloud • Security • Software • Cybersecurity • Automation
Provide technical direction for GitLab Dedicated, a managed single-tenant SaaS platform. Lead architecture and transformation across resilience, failover, tenant orchestration, change management, automation, and platform integrations. Identify systemic reliability and scalability risks, establish reusable platform patterns, strengthen service ownership, and guide cross-team technical decisions. Mentor senior engineers and advance engineering excellence across the organization.
Top Skills: Cloud InfrastructureDevsecopsDistributed SystemsGoInfrastructure As CodeObservabilityPythonRuby
Yesterday
Remote
Canada
Senior level
Senior level
Professional Services • Consulting
Own and improve CI/CD, developer tooling, agent harnesses, Kubernetes infrastructure, observability, cost management, and incident-response practices. Build platform capabilities that reduce engineering toil and improve delivery speed for conversational AI products. Participate in an on-call rotation for core infrastructure and help shape reliable, scalable systems across a fully remote engineering organization.
Top Skills: AlertingCi/CdHelmIncident ManagementKubernetesLlmsLogsMetricsMonitoringNode.jsObservabilityPythonTerraformTracingTypescript
2 Days Ago
In-Office or Remote
Toronto, ON, CAN
Mid level
Mid level
Software • PropTech
Responds to production incidents across AWS and Kubernetes, diagnosing and remediating infrastructure, networking, database, cache, and deployment issues. Builds monitoring, observability, synthetic and load tests, SLOs, and reliability improvements. Maintains Terraform, Kubernetes, CI/CD, IAM, autoscaling, and security configurations; leads postmortems and incident follow-up. Participates in on-call operations and collaborates with engineering teams to improve application performance and reliability.
Top Skills: .NetAmazon CloudwatchAmazon Ec2Amazon EksAmazon LinuxAmazon RdsAmazon S3AnsibleApplication Load BalancerAWSAzureBashCloudflareGithub ActionsGitlab CiGoogle Cloud PlatformGrafanaJavaScriptKubernetesLokiMemcachedMimirMySQLNetwork Load BalancerNginxOpentelemetryPagerdutyPHPPostgresPrometheusPythonRedisRuby On RailsRustTempoTerraformTraefikTypescriptUbuntu

What you need to know about the Toronto Tech Scene

Although home to some of the biggest names in tech, including Google, Microsoft and Amazon, Toronto has established itself as one of the largest startup ecosystems in the world. And with over 2,000 startups — more than 30 percent of the country's total startups — Toronto continues to attract new businesses. Be it helping entrepreneurs manage their finances, simplifying business operations by automating payroll or assisting pharmaceutical companies in launching new drugs, the city's tech scene is just getting started.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account