Shakudo Logo

Shakudo

Infrastructure Engineer

Reposted 18 Days Ago
Be an Early Applicant
In-Office
Toronto, ON, CAN
Senior level
In-Office
Toronto, ON, CAN
Senior level
Operate and maintain production infrastructure and internal services, including DGX/GPU clusters, bare-metal Kubernetes servers, CI/CD pipelines, and the customer-facing AI Gateway. Drive reliability, security hardening, observability, and DevOps best practices; contribute to product hardening and roadmap. Ensure uptime for internal and customer-facing systems.
The summary above was generated by AI
At Shakudo, we're building the world's first operating system for data and AI. We use the term "operating system" in the truest sense: just like iOS, Windows, or Linux, Shakudo's end-to-end OS provides ever-evolving, fully automated, best-in-class open-source components tailored to each business's unique needs.
 
We are seeking an Infrastructure Engineer to join our Business Automation team to own and operate the internal systems, infrastructure, and AI Gateway product that power Shakudo at scale. This is a hands-on role for someone who thrives on keeping production systems reliable, secure, and fast. You will be responsible for everything from physical servers and DGX machines to CI/CD pipelines and customer-facing AI Gateway infrastructure. You will also contribute directly to product hardening, security, and DevOps practices across the platform.
 
At Shakudo, our culture is proactive, collaborative, and supportive — we succeed together by building strong partnerships and solving complex challenges. We expect high ownership: you will be hands-on, driving outcomes directly rather than delegating or waiting for direction. Individual contribution matters here — your work will have a visible, measurable impact on the company's operations and product.

Key Responsibilities

  • Maintain and operate internal services for the rest of the Shakudo employees, including proprietary applications for sales and ETL pipelines
  • Maintain and operate DGX machines that host LLMs for the team's use
  • Maintain and operate Shakudo's product for Shakudo's internal use, and contribute to product hardening, security, and DevOps practices
  • Maintain and operate physical servers for Kubernetes clusters and ensure uptime
  • Create CI/CD pipelines for internal deployments
  • Maintain and operate the AI Gateway product for customers, ensure uptime, and contribute to product roadmap

Qualifications

  • 8+ years of experience across software, data, platform, or AI engineering roles
  • 5+ years of strong experience with Kubernetes cluster operation and DevOps, and bare-metal server operations
  • Experience operating production infrastructure at scale, including physical servers, GPU clusters, and CI/CD systems
  • Strong background in security hardening, observability, and reliability engineering
  • Proficiency in Rust is preferred
  • Experience with AI/ML infrastructure, including LLM hosting and inference serving is preferred 

Why Shakudo Stands Out

    Work with cutting-edge technologies in machine learning and high-performance computing. Contribute to a platform that transforms how organizations leverage data and AI. Join a dynamic team that values innovation, efficiency, and diversity.
     
    Shakudo offers a high-impact package: competitive salary, meaningful equity so you share in the upside of transformational technology, and comprehensive health benefits that have you fully covered. We provide a flexible vacation policy—because building transformational technology requires supporting the people who build it. More importantly, you'll work on technology that matters.
     
    This role is based onsite in Toronto to support the high security requirements of our clients and enable effective collaboration. We have a welcoming office environment with a very focused and passionate team, doing meaningful, impactful work together.
     
    Shakudo is an equal opportunity employer and encourages candidates of all backgrounds to apply. We foster diversity and inclusivity and welcome applications from a broad range of backgrounds and experiences.

HQ

Shakudo Toronto, Ontario, CAN Office

21 Carscadden Dr, Toronto, Ontario, Canada, M2R 2A6

Similar Jobs

4 Days Ago
In-Office or Remote
CA
Senior level
Senior level
Artificial Intelligence • Marketing Tech • Software • Generative AI
Own security for cloud and ML infrastructure, including tenant isolation, GPU cluster security, model protection, vulnerability management, network segmentation, monitoring, access governance, and credential management. Lead infrastructure security projects and establish technical standards and guardrails in partnership with Security, DevOps, Engineering, IT, and ML teams. The role also supports incident response, CSPM, sandboxing, container isolation, and SOC 2 readiness.
Top Skills: AWSKubernetesTerraform
6 Days Ago
Hybrid
Toronto, ON, CAN
Senior level
Senior level
Artificial Intelligence • Hardware • Software • Semiconductor
Designs and operates secure, scalable cloud infrastructure and identity platforms across AWS and data centers. Builds IAM, authentication, authorization, SSO, MFA, provisioning, secrets, and key management solutions. Develops Terraform-based automation using Python or Go, supports Kubernetes services, implements Zero Trust controls, and improves reliability, observability, and developer experience. The role includes architecture reviews, cross-team technical leadership, mentoring, production support, incident response, and on-call participation.
Top Skills: AWSAzureGCPGoIamIgaKubernetesMfaMicrosoft Entra IdOauth2OidcOktaPythonRbacSAMLScimSsoTerraformZero Trust
6 Days Ago
In-Office
Toronto, ON, CAN
Senior level
Senior level
Information Technology • Consulting • Cybersecurity
Designs and leads customer infrastructure solutions within professional services. Responsibilities include architectural planning, deployment documentation, deployment reviews, mentoring engineers, customer and project manager collaboration, troubleshooting, technical reviews, and maintaining vendor certifications. The role requires extensive experience with Dell EMC storage, virtualization, hyper-converged systems, networking, backup technologies, data centers, technical documentation, and onsite hardware installation. Travel up to 45% and after-hours support are required.
Top Skills: Azure StackBashCisco Tor NetworkingDell Data DomainDell Emc Powerscale/IsilonDell Emc PowerstoreDell Emc PowervaultDell Emc UnityDell Powerprotect Data ManagerDhcpDnsFiddlerHyper-VLayer 2 NetworkingLayer 3 NetworkingNetappNetcatNutanixPerlPowershellProxmoxPythonRoutingRp4VmVeeamVlansVmware VsanVmware VsphereVmware VxrailVmware/BroadcomVpnVsphere ReplicationWindows SysinternalsWireshark

What you need to know about the Toronto Tech Scene

Although home to some of the biggest names in tech, including Google, Microsoft and Amazon, Toronto has established itself as one of the largest startup ecosystems in the world. And with over 2,000 startups — more than 30 percent of the country's total startups — Toronto continues to attract new businesses. Be it helping entrepreneurs manage their finances, simplifying business operations by automating payroll or assisting pharmaceutical companies in launching new drugs, the city's tech scene is just getting started.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account