Fluidstack Logo

Fluidstack

Reliability Engineer, R&D

Reposted One Month Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in CA
Mid level
In-Office or Remote
Hiring Remotely in CA
Mid level
Lead reliability engineering for a reference design: build and validate availability and RAM models, run cross-discipline FMEAs, quantify failure rates and redundancy trade-offs, and close the loop by feeding fleet field-failure data back into designs to improve maintainability and uptime.
The summary above was generated by AI
About Fluidstack

We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it.

We're singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI.


We hire people who care deeply about this problem space. If that is you, please apply!

How We Operate
  • Be a barrel. Full autonomy. Own things end to end, take on scope without being asked, no permission required to operate outside your core role.

  • Insane urgency. We drive everything forward as fast as possible.

  • Reason from first principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.

  • Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.

  • Build something that actually matters. If you're going to spend your time, spend it on something that matters to the world.

Role Scope
  • Own reliability engineering for the reference design: availability modeled, weak points found, and fixes engineered before deployment.

  • Build the RAM models: failure rates, redundancy, and maintainability quantified per configuration.

  • Run FMEAs across disciplines: failure modes cataloged and designed out with the engineering teams.

  • Close the loop with the fleet: field failures fed back into models and design changes.

What We're Looking For
  • The below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly, tell us where you would.

  • You've done reliability engineering for infrastructure, energy, or complex hardware.

  • You've built availability models decision-makers used.

  • You've led cross-discipline FMEAs that changed designs.

  • You mine field data for the truth about failure rates.

  • You argue redundancy trade-offs in dollars and nines.

  • Bonus: Data center topologies. RAM modeling software. Weibull analysis. Maintenance strategy design.

We are committed to pay equity and transparency.

Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

You will receive a confirmation email once your application has successfully been accepted. If there is an error with your submission and you did not receive a confirmation email, please email [email protected] with your resume/CV, the role you've applied for, and the date you submitted your application-- someone from our recruiting team will be in touch.

Similar Jobs

Yesterday
Easy Apply
Remote
Canada
Easy Apply
Entry level
Entry level
Big Data • Fintech • Mobile • Payments • Financial Services
Build and improve reliable, scalable frontend experiences for Affirm’s pre-checkout products, including Affirm.js, promotional messaging, experimentation, and developer-facing tooling. Use browser debugging, monitoring, metrics, and testing tools to maintain accessible, performant cross-browser experiences. Collaborate with engineers and stakeholders, contribute to code reviews, troubleshoot operational issues, and write clear, well-tested, extensible code.
Top Skills: JavaScriptReactTypescriptVue
Yesterday
Remote or Hybrid
Toronto, ON, CAN
Senior level
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Design, build, deploy, and operate production optimization applications using machine learning, advanced algorithms, APIs, data pipelines, and evaluation systems. Own system correctness, performance, testing, reliability, observability, and customer outcomes. Direct AI coding agents through precise specifications and review their implementations. Collaborate with product, design, engineering, and customers to define requirements, communicate risks, and influence the product roadmap.
Top Skills: APIsAutomated TestingConcurrencyGraphQLJavaJavaScriptKubernetesMathematical OptimizationMultiprocessingObservabilityOperations ResearchParallelismPython
Yesterday
Easy Apply
Remote
Canada
Easy Apply
Senior level
Senior level
Cloud • Security • Software • Cybersecurity • Automation
Lead GitLab’s company-wide strategy for autonomous, agentic software development workflows. Identify and validate AI opportunities across the SDLC, build prototypes and reference architectures, establish evaluation frameworks and safety controls, and drive internal adoption. Partner across Engineering, Product, Infrastructure, Security, Architecture, and Data and ML teams to scale reliable, observable, compliant agentic capabilities into customer-facing products. Mentor senior technical leaders and represent GitLab externally on AI-assisted development and productivity measurement.
Top Skills: Agentic FrameworksAIAutonomous WorkflowsCi/CdData GovernanceDevsecopsDistributed SystemsGitlabLarge Language ModelsMachine LearningMulti-Tenant SystemsObservabilitySecurity And ComplianceSre

What you need to know about the Toronto Tech Scene

Although home to some of the biggest names in tech, including Google, Microsoft and Amazon, Toronto has established itself as one of the largest startup ecosystems in the world. And with over 2,000 startups — more than 30 percent of the country's total startups — Toronto continues to attract new businesses. Be it helping entrepreneurs manage their finances, simplifying business operations by automating payroll or assisting pharmaceutical companies in launching new drugs, the city's tech scene is just getting started.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account