Blackpoint Cyber Logo

Blackpoint Cyber

Sr. Site Reliability Engineer

Posted One Month Ago
Remote
Hiring Remotely in Canada
Senior level
Remote
Hiring Remotely in Canada
Senior level
Designs and maintains scalable cloud and on-premise infrastructure, Infrastructure as Code, Kubernetes environments, CI/CD pipelines, data streaming, caching, observability, and incident response systems. The role owns AWS reliability and cost efficiency, supports progressive deployments, troubleshoots complex production issues, partners with software teams, and drives automation and continuous improvement across platform operations.
The summary above was generated by AI

Blackpoint Cyber is the leading provider of world-class cybersecurity threat hunting, detection and remediation technology. Founded by former National Security Agency (NSA) cyber operations experts who applied their learnings to bring national security-grade technology solutions to commercial customers around the world, Blackpoint Cyber is in hyper-growth mode,  fueled by a recent $190m series C round. 

SUMMARY

We're hiring a Senior Site Reliability Engineer to design, implement, and maintain our cloud and on-premise infrastructure and CI/CD pipelines, with a focus on automation, scalability, and performance. You'll work across cloud platform administration, container orchestration, data streaming, observability, and incident response — partnering with engineering teams to keep our systems reliable, secure, and efficient, and helping foster a culture of continuous improvement.

RESPONSIBILITIES

  • Design, develop, and maintain highly scalable infrastructure using Infrastructure as Code (Terraform and Terragrunt) for automated cloud resource provisioning and orchestration.

  • Own and optimize our AWS cloud environment, ensuring cost efficiency, security best practices, and high-availability standards.

  • Manage and optimize Kubernetes cluster environments (Helm, ArgoCD, Istio, Kustomize) to support continuous delivery and infrastructure-as-code practices.

  • Administer and scale data streaming infrastructure (Confluent Cloud, Apache Kafka) to support enterprise-level data processing.

  • Deploy, configure, and maintain Redis for caching and real-time data processing.

  • Implement and maintain monitoring, alerting, and incident response frameworks (Prometheus, Grafana, Alert Manager, Grafana CloudOpsGenie/PagerDuty) to ensure system reliability and performance.

  • Facilitate controlled feature deployments and progressive rollouts through LaunchDarkly/PostHog.

  • Partner with software development teams to ensure seamless integration of new services, applications, and features into existing infrastructure.

  • Diagnose and resolve complex system-level issues, implementing solutions that maintain high performance and maximize uptime.

  • Drive continuous improvement of automation tooling, operational processes, and engineering methodologies to enhance scalability, reliability, and maintainability.

  • Stay current on emerging SRE trends and tools, andtools and help the team adopt relevant industry advancements and best practices.

REQUIREMENTS

  • 5+ years of experience in a Senior Site Reliability Engineer role or equivalent, with substantial emphasis on cloud infrastructure management and automation.

  • Expertise in Infrastructure as Code (Terraform, Terragrunt) for enterprise-scale deployments.

  • Comprehensive knowledge of AWS, including designing, implementing, and maintaining secure, scalable, resilient cloud architectures.

  • Extensive hands-on experience with distributed data streaming (Confluent Cloud, Apache Kafka).

  • Proven experience with Redis for caching and Amazon RDS for relational database management.

  • Experience with enterprise search and analytics platforms (OpenSearch, Elasticsearch, ChaosSearch).

  • Proficiency designing and implementing monitoring/alerting infrastructure (Prometheus, Grafana, Alert Manager, Grafana Cloud, OpsGenie/PagerDuty).

  • Practical experience with feature flag systems (LaunchDarkly/PostHog) for controlled release management.

  • Extensive experience administering production-grade Kubernetes (Helm, ArgoCD, Istio); working knowledge of Kustomize.

  • Strong problem-solving skills, with the ability to troubleshoot complex systems in production.

  • Strong communication and collaboration skills, with experience working in Agile environments.

NICE TO HAVE

  • Experience with Terragrunt to manage Terraform across multiple environments.

  • Extensive hands-on experience with distributed data streaming (Kafka).

  • Multi-cloud experience (Google Cloud Platform, Microsoft Azure).

  • Understanding of security frameworks and compliance standards for cloud-native/containerized environments.

  • Serverless computing and CI/CD pipeline experience (Jenkins, GitHub Actions).

  • Software development proficiency in Node.js, Python, and/or Go.

Blackpoint Cyber welcomes and encourages applications from qualified individuals of all races, colors, religions, sex, sexual orientation, gender identity or expression, national origin, age, marital status, or any other legally protected status. We are committed to equality of opportunity in all aspects of employment.

For eligible employees in the US, Blackpoint offers competitive Health, Vision, Dental, and Life Insurance plans, a robust 401k plan, Discretionary Time Off, and other minor perks. International employees receive competitive benefits in accordance with local market standards and applicable country requirements.

Blackpoint believes all employees should share in the company’s success – equity participation is available to employees globally, with program details varying by location and employment structure.

Similar Jobs

Yesterday
Remote
Canada
Senior level
Senior level
Software
Build and operate scalable, resilient, distributed services across AWS and Azure. The Senior Site Reliability Developer will automate infrastructure, develop platforms and frameworks, establish monitoring and incident response practices, define runbooks, reduce operational toil, lead delivery projects, mentor engineers, and participate in technical interviews and on-call rotations. The role requires expertise in cloud infrastructure, infrastructure as code, CI/CD, Kubernetes, observability, telemetry, and SRE practices including SLOs, SLIs, and error budgets.
Top Skills: Amazon CloudwatchAmazon EcrAmazon EcsAmazon EksAmazon KinesisAmazon RdsAmazon RedshiftAmazon SqsAnsibleArtifactoryAuth0AWSAzureAzure Container AppsAzure DevopsAzure MonitorCloudamqpDockerDropwizardElasticsearchGitGithub ActionsGitlabHibernateJavaJavaScriptJenkinsKubernetesMongoDBMySQLObserve Inc.OpentelemetryPrometheusPythonRabbitMQReactRedshift SpectrumReduxSpring BootTemporalTerraformTypescript
Yesterday
Remote
Canada
Senior level
Senior level
Software
Operate and maintain highly available, secure, containerized SaaS applications across AWS and Azure. Responsibilities include rotating 24x7 on-call coverage, observability, incident and security response, disaster recovery, Terraform infrastructure-as-code, CI/CD automation, cloud integration, performance optimization, and developing AI agents to automate SRE and DevSecOps workflows.
Top Skills: Amazon Web ServicesCheckovClaude CodeDatadogDockerDynatraceGithub ActionsGitlab CiGoJavaScriptKubernetesAzureNew RelicPrisma CloudPythonTerraformWiz
5 Days Ago
Remote
Canada
Senior level
Senior level
Cloud • Security • Software • Generative AI
Own and improve Elastic’s observability infrastructure across hosted cloud deployments. Responsibilities include Terraform-based infrastructure delivery, Python and Go development, production operations, incident response, on-call participation, RCA and postmortem writing, code and design reviews, mentoring, and improving operational documentation and processes. The role also involves operating Linux and containerized workloads, delivering complex projects independently, and maintaining secure, reliable platform infrastructure.
Top Skills: AnsibleArgocdBeatsElastic Cloud Enterprise (Ece)Elastic Cloud Hosted (Ech)Elastic Cloud On Kubernetes (Eck)ElasticsearchGoHelmKibanaKubernetesKyvernoLinuxLogstashPuppetPythonTeleportTerraformVault

What you need to know about the Toronto Tech Scene

Although home to some of the biggest names in tech, including Google, Microsoft and Amazon, Toronto has established itself as one of the largest startup ecosystems in the world. And with over 2,000 startups — more than 30 percent of the country's total startups — Toronto continues to attract new businesses. Be it helping entrepreneurs manage their finances, simplifying business operations by automating payroll or assisting pharmaceutical companies in launching new drugs, the city's tech scene is just getting started.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account