Affirm Logo

Affirm

Manager, Software Engineering (Resilience Engineering)

Posted 22 Days Ago
Be an Early Applicant
Easy Apply
Remote
Hiring Remotely in Canada
Entry level
Easy Apply
Remote
Hiring Remotely in Canada
Entry level
Leads Affirm’s Resilience Engineering team, defining strategy and managing engineers who build platforms for safe production load testing and chaos engineering. Owns fault-injection and experimentation systems, safeguards, observability, rollback mechanisms, monitoring, and incident response practices. Partners cross-functionally to identify systemic weaknesses, improve reliability, and embed resilience validation into engineering workflows. Requires leadership experience in reliability, infrastructure, or distributed systems and hands-on expertise with production testing, chaos engineering, cloud-native environments, and distributed-system failure modes.
The summary above was generated by AI

At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most.

We are seeking a seasoned Engineering Manager to lead our Resilience Engineering team. This role is critical in ensuring the safety and reliability of our production systems through proactive validation techniques, including production load testing and chaos engineering.

You will lead the development of systems and practices that allow engineers to safely test system behavior under stress and failure conditions in production, ensuring issues are discovered and mitigated before they impact real users.


What you’ll doLeadership & Strategy
  • Define and drive the vision for resilience engineering at Affirm, with a focus on production load testing and chaos engineering as first-class engineering practices.
  • Lead and mentor a team of engineers building platforms and tooling for safe production experimentation.
  • Partner with infrastructure, product, and security leadership to embed resilience validation into the software development lifecycle.
  • Establish best practices for safely testing system limits and failure scenarios in production.

Systems & Operations
  • Own the design and evolution of platforms that enable safe, controlled production load testing and fault injection.
  • Ensure strong safeguards are in place, including isolation boundaries, approval workflows, and automated rollback mechanisms to protect real users.
  • Build systems that provide end-to-end observability, traceability, and auditability for all resilience experiments.
  • Drive reliability improvements by systematically identifying weaknesses through load testing and chaos experiments.
  • Establish monitoring, alerting, and incident response practices tailored to proactive resilience validation.
Collaboration & Enablement
  • Work closely with engineering teams to design and execute production load tests and chaos experiments safely.
  • Partner with infrastructure teams to build guardrails around tests and experimentations.
  • Enable teams to adopt resilience practices by providing reusable tooling, frameworks, and standardized workflows.
  • Identify systemic weaknesses and lead cross-functional efforts to improve reliability and fault tolerance.
  • Evangelize a culture of “test failure before failure tests you” across the organization.

What we look for
  • Proven experience leading engineering teams in reliability, infrastructure, or distributed systems.
  • Hands-on experience with production load testing, chaos engineering, or large-scale system validation.
  • Experience with leveraging a chaos engineering vendor such as Gremlin, Harness, or something similar.
  • Strong understanding of failure modes in distributed systems, including latency, partial failure, and cascading outages.
  • Experience building or operating systems with strong safety guarantees (isolation, rate limiting, guardrails, auditability).
  • Familiarity with cloud-native environments (AWS, Kubernetes) and observability tooling.
  • Strong programming background (e.g., Python, Kotlin, Java, or similar).
  • Excellent problem-solving skills and the ability to balance long-term resilience investments with immediate business needs.
  • Strong communication and leadership skills, with a track record of influencing engineering practices across teams.

This posting is for an existing vacancy.

Pay Grade - P
Equity Grade - 7
Employees new to Affirm typically come in at the start of the pay range. Affirm focuses on providing a simple and transparent pay structure which is based on a variety of factors, including location, experience and job-related skills. 
Base pay is part of a total compensation package that may include monthly stipends for health, wellness and tech spending, and benefits (including 100% subsidized medical coverage, dental and vision for you and your dependents). In addition, the employees may be eligible for equity rewards offered by Affirm Holdings, Inc. (parent company).
CAN base pay range per year: 181,000 - 241,000

#LI-Remote

 

Remote-first with flexibility built in
Affirm is proud to be a remote-first company. Most roles can be done from almost anywhere within the country of employment. Some positions may occasionally require in-person work at an Affirm office, and a few are office-based due to the nature of the work. All new hires will be invited to attend an in-person onboarding experience.

Benefits designed for you
Our benefits reflect our commitment to care, transparency, and flexibility. Here are a few highlights:

  • Health coverage at no cost: We cover 100% of premiums for employees and their dependents.
  • Spending stipends: Monthly stipends support your tech setup, and the ability to choose health and wellness options that are right for you.
  • Time off to recharge: Flexible time off and generous holiday calendars help you rest when you need to.
  • Own a piece of what you build: Our employee stock purchase plan (ESPP) lets you buy Affirm stock at a discount.

We’re committed to providing an inclusive interview process, including accommodations for candidates with disabilities. If you need support, we’re happy to help.

For positions based in San Francisco or Los Angeles: Affirm considers qualified applicants with arrest and conviction records, as required by law.

By clicking "Submit Application," you acknowledge that you have read Affirm's Global Candidate Privacy Notice and consent to the use of your personal information as described.

Affirm Toronto, Ontario, CAN Office

Toronto, ON, Canada

Similar Jobs at Affirm

Yesterday
Easy Apply
Remote
Canada
Easy Apply
Senior level
Senior level
Big Data • Fintech • Mobile • Payments • Financial Services
Own and deliver high-availability authentication, verification, and fraud experiences across web, backend, and mobile interfaces. Lead quarterly goals, technical planning, system delivery, code quality, operational availability, and on-call support. Collaborate with product, design, and analytics stakeholders, resolve complex technical and business issues, mentor engineers, and promote engineering standards. The role requires expertise in React or Vue, JavaScript or TypeScript, large codebases, scalable architecture, and effective communication.
Top Skills: JavaScriptReactTypescriptVue
2 Days Ago
Easy Apply
Remote
Canada
Easy Apply
Senior level
Senior level
Big Data • Fintech • Mobile • Payments • Financial Services
Design, develop, launch, and maintain scalable backend systems and APIs used by partners and merchants. Collaborate with engineering teams and stakeholders, balance delivery speed with system quality and reliability, navigate large codebases, debug software, conduct code reviews, and write clear, tested, extensible code using distributed-systems technologies.
Top Skills: AWSKotlinKubernetesMySQLPython
7 Days Ago
Easy Apply
Remote
Canada
Easy Apply
Junior
Junior
Big Data • Fintech • Mobile • Payments • Financial Services
Build and operate backend systems supporting post-transaction card account accuracy, issue resolution, and customer communications. Break down projects, deliver work in phases, collaborate with product and cross-functional teams, monitor system metrics, support availability and on-call operations, and contribute to code reviews and interviews. The role requires backend development experience, distributed systems knowledge, and proficiency in Python or Kotlin, AWS, MySQL, and Kubernetes.
Top Skills: AWSKotlinKubernetesMySQLPython

What you need to know about the Toronto Tech Scene

Although home to some of the biggest names in tech, including Google, Microsoft and Amazon, Toronto has established itself as one of the largest startup ecosystems in the world. And with over 2,000 startups — more than 30 percent of the country's total startups — Toronto continues to attract new businesses. Be it helping entrepreneurs manage their finances, simplifying business operations by automating payroll or assisting pharmaceutical companies in launching new drugs, the city's tech scene is just getting started.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account