AuthZed Logo

AuthZed

Sr. Site Reliability Engineer

Reposted One Month Ago
Remote
Hiring Remotely in Canada
Senior level
Remote
Hiring Remotely in Canada
Senior level
As a Site Reliability Engineer, you will design, implement, and maintain scalable infrastructure, ensure system reliability, automate processes, and collaborate with engineering teams.
The summary above was generated by AI
About AuthZed:

We are the creators and maintainers of SpiceDB and the authorization infrastructure that companies around the world depend on to keep their engineering teams focused on what matters most - their own product.

We are a Series A company, fixing broken access control with products that eliminate complex permission management while delivering enterprise-scale performance and consistent access control.

AuthZed is a fully remote company with employees across the US, Canada, and Europe. We’re a hardworking and close-knit group with a software-driven culture (yep, even our GTM team understands and loves this technology)! We bring integrity to all our interactions, fostering confidence in decision making - trusting and respecting each voice on our team, every day.

Company Values:
  • Agency: Everyone should have the capability, freedom, and confidence to bring about changes to our business and product. Organizational processes exist to clearly define our goals, but not restrict how progress is made.

  • Collaboration: Success is defined in various dimensions and no single person can be an expert in all of them. Without valuing the opinions of others, finding compromises, and sharing mutual trust and respect, you cannot arrive at the best possible solution.

  • Open-mindedness: Without asking questions, testing assumptions, and questioning our pre-existing biases we risk operating within an echo-chamber. We celebrate the representation of diverse perspectives and backgrounds as a catalyst for creating an inclusive work environment that everyone can appreciate.

About the Role:

As a Site Reliability Engineer, you will play a critical role in ensuring the reliability, availability, and performance of our systems. You will be responsible for designing, implementing, and maintaining scalable infrastructure solutions to support our growing customer base. This is an exciting opportunity to work in a fast-paced environment and contribute to the success of a company bringing a Google-inspired authorization system to companies around the globe.

What you’ll own:
  • Design, implement, and maintain highly available and scalable infrastructure solutions for our projects, products, and customers.

  • Monitor and analyze system performance, identifying and resolving bottlenecks and issues to ensure optimal performance and reliability.

  • Automate infrastructure deployment and configuration management processes.

  • Continuously improve system reliability, security, and efficiency through proactive monitoring, capacity planning, and performance tuning.

  • Troubleshoot and resolve complex infrastructure and application issues in production and test environments.

  • Collaborate with software engineering teams to design and implement systems that are resilient, scalable, and secure.

  • Participate in on-call rotation and respond to production incidents in a timely manner.

  • Document system configurations, troubleshooting procedures, and operational guidelines.

What you bring:
  • Proven experience as a Site Reliability Engineer or in a similar role.

  • Strong understanding of networking, operating systems, and cloud infrastructure.

  • Experience with Site Reliability Engineering, System Design, and Distributed Computing.

  • Experience in various programming languages — we currently have SDKs for NodeJS, Java, Python, Ruby, and Go.

  • Experience with containerization technologies such as Docker and Kubernetes.

  • Knowledge of infrastructure-as-code tools like Terraform and Pulumi.

  • Familiarity with monitoring and logging tools (e.g., Prometheus, Grafana, ELK stack).

  • Experience with lower-level implementation details of relational databases (bonus if you have have experience with distributed SQL databased like Google Cloud Spanner or CockroachDB).

  • Experience working with Git and GitHub.

  • Experience with continuous integration and deployment systems.

  • Strong problem-solving and troubleshooting skills.

  • Excellent communication and collaboration abilities.

Extra shine:
  • Experience with Authorization systems.

Life at AuthZed:
  • Opportunity to work with cutting-edge technology in a rapidly growing sector.

  • A supported environment where your ideas lead to real impact.

  • Competitive salary based on experience.

  • Stock options at an early-stage startup.

  • Comprehensive benefits including healthcare (US-based) and other insurance.

  • A full remote and flexible schedule to accommodate different timezones

  • Twice-yearly travel for team offsites focused on team bonding, collaboration, and having fun!

Similar Jobs

5 Days Ago
Remote
Canada
Senior level
Senior level
Software
Build and operate scalable, resilient, distributed services across AWS and Azure. The Senior Site Reliability Developer will automate infrastructure, develop platforms and frameworks, establish monitoring and incident response practices, define runbooks, reduce operational toil, lead delivery projects, mentor engineers, and participate in technical interviews and on-call rotations. The role requires expertise in cloud infrastructure, infrastructure as code, CI/CD, Kubernetes, observability, telemetry, and SRE practices including SLOs, SLIs, and error budgets.
Top Skills: Amazon CloudwatchAmazon EcrAmazon EcsAmazon EksAmazon KinesisAmazon RdsAmazon RedshiftAmazon SqsAnsibleArtifactoryAuth0AWSAzureAzure Container AppsAzure DevopsAzure MonitorCloudamqpDockerDropwizardElasticsearchGitGithub ActionsGitlabHibernateJavaJavaScriptJenkinsKubernetesMongoDBMySQLObserve Inc.OpentelemetryPrometheusPythonRabbitMQReactRedshift SpectrumReduxSpring BootTemporalTerraformTypescript
5 Days Ago
Remote
Canada
Senior level
Senior level
Software
Operate and maintain highly available, secure, containerized SaaS applications across AWS and Azure. Responsibilities include rotating 24x7 on-call coverage, observability, incident and security response, disaster recovery, Terraform infrastructure-as-code, CI/CD automation, cloud integration, performance optimization, and developing AI agents to automate SRE and DevSecOps workflows.
Top Skills: Amazon Web ServicesCheckovClaude CodeDatadogDockerDynatraceGithub ActionsGitlab CiGoJavaScriptKubernetesAzureNew RelicPrisma CloudPythonTerraformWiz
6 Days Ago
Remote
Canada
Senior level
Senior level
Artificial Intelligence • Information Technology • Software • Automation
Designs and maintains Camunda’s Kubernetes-based multi-cloud infrastructure, improving scalability, reliability, monitoring, observability, and automation. Owns systems end-to-end, participates in incident response and on-call rotations, creates runbooks, collaborates with product and engineering teams, and mentors less experienced engineers. The role requires strong Kubernetes, infrastructure-as-code, monitoring, incident response, and automation expertise, with cloud, GitOps, programming, and SLO experience preferred.
Top Skills: Amazon EksArgocdAWSGitopsGoGoogle Cloud PlatformGoogle Kubernetes EngineGrafanaInfrastructure As CodeKubernetesMonitoringMulti-Cloud InfrastructureObservabilityPrometheusPythonTerraform

What you need to know about the Toronto Tech Scene

Although home to some of the biggest names in tech, including Google, Microsoft and Amazon, Toronto has established itself as one of the largest startup ecosystems in the world. And with over 2,000 startups — more than 30 percent of the country's total startups — Toronto continues to attract new businesses. Be it helping entrepreneurs manage their finances, simplifying business operations by automating payroll or assisting pharmaceutical companies in launching new drugs, the city's tech scene is just getting started.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account