Senior DevOps / Site Reliability Engineer to build, operate, monitor, and scale an enterprise AI platform on Azure. Responsibilities include Azure infrastructure and Kubernetes deployment, IaC and CI/CD automation, advanced dashboards and observability, monitoring/alerting, capacity planning and automated scaling, troubleshooting, and collaborating with developers and AI engineers. Hybrid role based in the GTA with occasional client travel.
Job Title:- Senior DevOps Engineer
Work Type:- Full-Time
Work Mode:- Hybrid
Tru Inc is seeking a Senior DevOps / Site Reliability Engineer . We have one (1) permanent full-time position available for an immediate start.
About Us:-
At Tru, we are more than just a technology agency—we are strategic partners committed to driving digital transformation for enterprise-level clients. As experts in digital experience platforms, e-commerce engines, and innovative solutions, we help businesses thrive in the digital era.
Position Overview :-
We are looking for a Senior DevOps / Site Reliability Engineer to help build, operate, monitor and scale our new enterprise AI platform. This role will be responsible for the DevOps and reliability capabilities supporting the platform across all environments including Development, QA and Production. The environment is currently focused and manageable consisting of approximately 10–30 containers but is expected to grow as the AI platform expands in 2027 and beyond.
This is an opportunity to work with a collaborative team on a new platform where the focus is not simply deploying infrastructure. The most important part of the role will be creating intelligent dashboards, monitoring, alerting, automated scaling and AI-assisted platform management. The platform is hosted primarily in Microsoft Azure with limited exposure to AWS.
Location:-
This role is hybrid, providing some flexibility to work from home. However, candidates must reside in the Greater Toronto Area (GTA) and be prepared to attend in-person meetings as required. Occasional travel to client sites or workshops may also be necessary.
Key Responsibilities:-
Azure Infrastructure and Platform Deployment:-
- Design, deploy, configure and maintain infrastructure within Microsoft Azure.
- Deploy and manage virtual machines, containers, Kubernetes clusters, networking, storage and supporting platform services.
- Support infrastructure across Development, QA and Production environments.
- Establish repeatable and reliable deployment processes using Infrastructure as Code and CI/CD automation.
- Maintain secure, resilient and appropriately sized platform environments.
Kubernetes and Container Management:-
- Deploy, configure and operate containerized applications using Kubernetes.
- Manage container lifecycle, configuration, secrets, networking, storage and application dependencies.
- Monitor container and cluster health, resource consumption, capacity and performance.
- Troubleshoot deployment, networking, configuration and runtime issues.
- Establish appropriate standards for container deployment and Kubernetes operations.
Performance, Load Management, and Scaling:-
- Monitor platform demand, workload patterns, resource utilization and application performance.
- Configure horizontal and vertical scaling policies for containers and supporting infrastructure.
- Develop intelligent scaling approaches based on workload, queue depth, response time, resource utilization and business demand.
- Conduct capacity planning and identify potential performance bottlenecks before they affect production.
- Help introduce predictive or AI-assisted scaling and platform management capabilities.
Dashboards and Platform Visibility:-
- Design and build advanced operational dashboards using tools such as Grafana, Kibana, Azure Monitor, Application Insights and similar technologies.
- Create clear executive, operational, application and infrastructure views of platform health.
- Build dashboards covering availability, performance, capacity, errors, latency, traffic, container health, AI workloads and service dependencies.
- Establish meaningful service-level indicators, service-level objectives and reliability metrics.
- Continuously improve dashboards so that issues, trends and risks can be quickly identified.
- Advanced dashboard design and dashboard-building experience is a core requirement for this role.
Monitoring and Alerting:-
- Implement monitoring and alerting across infrastructure, applications, containers, integrations and AI platform services.
- Configure actionable alerts that identify real production risks while minimizing unnecessary alert noise.
- Establish thresholds, anomaly detection, health checks, synthetic monitoring and automated remediation where appropriate.
- Create operational runbooks and troubleshooting guidance.
- Work with development and architecture teams to improve platform observability.
Production Reliability and Support:-
- Support the stability, availability and operational readiness of the production AI platform.
- Investigate and resolve platform, deployment, infrastructure, monitoring and performance issues.
- Participate in root-cause analysis and implement preventative improvements.
- Ensure that production support processes, documentation and escalation paths are established before platform usage increases.
- Provide very light production support during 2026, with no regular after-hours support currently anticipated.
- Help prepare the operating model for increased platform adoption and support requirements expected in 2027.
Requirements:-
- Strong professional experience in DevOps, Site Reliability Engineering, cloud infrastructure or platform engineering.
- Advanced hands-on experience with Microsoft Azure.
- Strong experience deploying and operating Kubernetes environments.
- Strong knowledge of containerization technologies such as Docker.
- Experience deploying and supporting containerized applications in Development, QA and Production environments.
- Advanced experience designing and building dashboards using Grafana, Kibana, Azure Monitor, Application Insights or comparable tools.
- Strong experience implementing monitoring, observability, logging, alerting and operational health checks.
- Experience managing application load, infrastructure capacity, performance and automated scaling.
- Experience with CI/CD pipelines and automated application deployment.
- Experience with Infrastructure as Code tools such as Terraform, Bicep or ARM templates.
- Strong troubleshooting skills across applications, containers, infrastructure, networking and cloud services.
- Ability to work independently while collaborating closely with developers, architects, AI engineers and platform stakeholders.
Preferred Experience:-
- Experience supporting AI, machine learning, data or high-compute platforms.
- Experience monitoring AI models, inference services, token usage, GPU workloads, API consumption, queues or model performance.
- Experience implementing automated remediation, predictive monitoring or AI-assisted platform operations.
- Familiarity with AWS services and cloud operations.
- Experience with Elasticsearch, Log Analytics, OpenTelemetry, Prometheus or similar observability technologies.
- Experience defining service-level indicators, service-level objectives and reliability standards.
- Experience with security, identity, secrets management and cloud governance within Azure.
What Makes This Role Different:-
This is not a large-scale high-pressure production support environment. The initial platform footprint is relatively focused with approximately 10–30 containers and very limited production support expected during 2026.
The role offers the opportunity to establish the platform correctly from the beginning, introduce modern DevOps and SRE practices and experiment with intelligent monitoring, automated scaling, advanced dashboards and AI-assisted platform management. As platform adoption increases the responsibilities and operational scope are expected to grow throughout 2027.
Benefits:-
Salary Range:- $100,000-$120,000
Working at Tru:-
At Tru, we put people first! We take pride in building a culture that stands out for its courage, entrepreneurial culture, diversity, and passion for people. We are proud to offer competitive salaries along with a 100% employer paid benefits package with a remote work- from- home arrangement. At Tru, we believe in creating an environment that is challenging, fun, and rewarding. We have regular team events to enhance our team spirit, including fun live gatherings that bring us together beyond work.
At Tru, we are a family and our embraced values at Tru are:
- You Talk, We Listen
- Integrity at Our Core
- Quality as Standard
- Delivered On Time
Join us and be part of a team that makes a tangible impact in the digital world!
Similar Jobs
Agency • Fashion • News + Entertainment • Sports
Owner of cloud infrastructure and reliability for a distributed microservices platform: design and maintain Azure infrastructure with Terraform, manage Kubernetes/Helm deployments, build CI/CD pipelines and observability, operate PostgreSQL at scale, lead incident response and postmortems, and support migration from legacy .NET/MSSQL to a cloud-native stack.
Top Skills:
.NetAksAWSAzureAzure AdAzure Blob StorageAzure DevopsAzure Key VaultDatadogDockerGithub ActionsGrafanaGrpcHashicorp VaultHelmKubernetesMssqlNext.JsOpentelemetryPostgresTerraform
Big Data • Cloud • Digital Media • Machine Learning • Mobile • Software • Industrial
Lead and evolve a cloud-native platform by improving CI/CD, IaC, reliability, security, and automation. Architect scalable hosting, drive DevOps best practices, manage AWS-based infrastructure, participate in on-call and incident response, and mentor cross-functional teams to increase platform resilience and delivery velocity.
Top Skills:
ArtifactoryAWSBashCloudFormationCloudwatchDockerDynamoDBDynatraceEc2EcsEksElasticsearchElbGitIamJavaJenkinsKubernetesLambdaLinuxMySQLPostgresPythonRdsS3SplunkSQLTerraformVpc
Aerospace • Automotive • Robotics
Own and operate infrastructure-as-code across cloud, on-prem, and bare-metal; build CI/CD pipelines; containerize and orchestrate services; maintain observability and reliability; participate in on-call and incident response; enable engineering teams with self-service tooling and secure deployment practices.
Top Skills:
AWSDockerGithub ActionsGrafanaJenkinsKubernetesLinuxPrometheusSaltstackTerraform
What you need to know about the Toronto Tech Scene
Although home to some of the biggest names in tech, including Google, Microsoft and Amazon, Toronto has established itself as one of the largest startup ecosystems in the world. And with over 2,000 startups — more than 30 percent of the country's total startups — Toronto continues to attract new businesses. Be it helping entrepreneurs manage their finances, simplifying business operations by automating payroll or assisting pharmaceutical companies in launching new drugs, the city's tech scene is just getting started.

.png)
