Tenstorrent Inc. Logo

Tenstorrent Inc.

Staff, Reliability Engineer

Reposted One Month Ago
Be an Early Applicant
In-Office
Toronto, ON, CAN
Senior level
In-Office
Toronto, ON, CAN
Senior level
Lead reliability strategy for next-generation AI computing systems. Drive architecture-to-production reliability, perform root-cause analyses, validate advanced cooling, coordinate cross-disciplinary teams, engage manufacturing partners, and mentor engineers to improve uptime and durability.
The summary above was generated by AI

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities.

Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy behind the next generation of AI computing systems. In this highly visible technical leadership role, you'll drive reliability from architecture through production, partnering across hardware, software, and manufacturing teams to build high-performance AI platforms that set the standard for uptime, durability, and quality. If you're passionate about solving complex engineering challenges and influencing products at scale, you'll have the opportunity to shape technology powering the future of AI.

This role is hybrid, based out of Toronto, Canada.

We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting.


Who You Are

  • You've spent 8+ years in reliability engineering, ideally in high-performance computing, AI hardware, or data center systems.
  • You're comfortable with the statistical side of the job, HALT, HASS, ALT, MTBF, Weibull analysis, and FMEA are all familiar territory.
  • You can work through a technical problem in a thermal lab and then explain the risks and trade-offs clearly to leadership.
  • You're good at bringing people together, mechanical, electrical, thermal, software, and System Dev & Compliance Validation teams, especially when timelines are tight.
  • You hold a Bachelor's or Master's in Mechanical Engineering, Electrical Engineering, Reliability Engineering, or a related field.

What We Need

  • Someone to set the reliability strategy for our next-generation AI computing systems, with MTBF modeling as a core piece.
  • A strong problem-solver who can lead root-cause investigations and drive fixes across engineering and the supply chain.
  • Someone who summarizes findings and feeds them back to the Systems Engineering design team, reliability as an ongoing loop, not a one-time check.
  • A close partner to System Dev & Compliance Validation, helping set hardware up for success ahead of formal testing and certification.
  • Willingness to travel to third-party test facilities for HALT/HASS, to preplan, oversee testing, resolve DUT issues, and assess design risk in person.

What You Will Learn

  • How to build a reliability strategy from scratch for hardware that's pushing the boundaries of AI computing.
  • How to build predictive models, including MTBF frameworks and accelerated life testing.
  • How to turn test findings into design improvements through close collaboration with Systems Engineering.
  • How reliability work sets the stage for validation and certification success.
  • What it takes to run hands-on testing at manufacturing and test partner sites, including in Taiwan.

Tenstorrent offers a highly competitive compensation package and benefits, and we are an equal opportunity employer.

This offer of employment is contingent upon the applicant being eligible to access U.S. export-controlled technology.  Due to U.S. export laws, including those codified in the U.S. Export Administration Regulations (EAR), the Company is required to ensure compliance with these laws when transferring technology to nationals of certain countries (such as EAR Country Groups D:1, E1, and E2).   These requirements apply to persons located in the U.S. and all countries outside the U.S.  As the position offered will have direct and/or indirect access to information, systems, or technologies subject to these laws, the offer may be contingent upon your citizenship/permanent residency status or ability to obtain prior license approval from the U.S. Commerce Department or applicable federal agency.  If employment is not possible due to U.S. export laws, any offer of employment will be rescinded.

HQ

Tenstorrent Inc. Toronto, Ontario, CAN Office

Toronto, ON, Canada

Similar Jobs

One Month Ago
Easy Apply
Hybrid
Toronto, ON, CAN
Easy Apply
Senior level
Senior level
Cloud • Mobile • Software
Lead company-wide technical strategy for software quality and reliability. Drive cross-team initiatives, establish architecture and engineering standards, build shared platform capabilities (release safety, automated validation, observability), improve testability and resilience, define reliability metrics, lead multi-team programs, and mentor engineers to raise organizational reliability practices.
Top Skills: Automated ValidationAWSCi/CdDeveloper PlatformsJavaObservabilityPerformance EngineeringRelease SafetyResilience EngineeringTypescript
10 Days Ago
Hybrid
Toronto, ON, CAN
Senior level
Senior level
AdTech • Software
Design and build Index Cloud, a globally distributed, multi-tenant compute platform across Kubernetes, bare-metal, and public cloud infrastructure. Own architectural decisions, infrastructure-as-code frameworks, platform APIs, SDKs, deployment systems, networking, storage, and distributed systems solutions. Drive technical strategy through RFCs and design reviews, improve developer productivity with self-service tooling, and mentor engineers across multiple teams.
Top Skills: AnsibleArgocdAWSCephDnsEksElkGCPGitopsGkeGoGrafanaHadoopHbaseKafkaKubernetesLinuxLoad BalancingLokiMimirNetworkingPrometheusPythonService DiscoverySparkTempoTerraformVault
10 Days Ago
Hybrid
Toronto, ON, CAN
Expert/Leader
Expert/Leader
AdTech • Software
Design and build Index Cloud, a globally distributed, multi-tenant platform operating across bare-metal and public cloud infrastructure. Responsibilities include architecting Kubernetes clusters, infrastructure-as-code frameworks, deployment systems, platform APIs, SDKs, and distributed systems. The role drives technical strategy through RFCs and design reviews, improves developer self-service, establishes security and tooling standards, mentors engineers, and collaborates across SRE, networking, security, operations, and software engineering teams.
Top Skills: Amazon EksAnsibleArgocdAWSCephDnsElkGCPGitopsGoGoogle GkeGrafanaHadoopHbaseKafkaKubernetesLinuxLokiMimirPrometheusPythonSparkTempoTerraformVault

What you need to know about the Toronto Tech Scene

Although home to some of the biggest names in tech, including Google, Microsoft and Amazon, Toronto has established itself as one of the largest startup ecosystems in the world. And with over 2,000 startups — more than 30 percent of the country's total startups — Toronto continues to attract new businesses. Be it helping entrepreneurs manage their finances, simplifying business operations by automating payroll or assisting pharmaceutical companies in launching new drugs, the city's tech scene is just getting started.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account