Design, orchestrate, and optimize large-scale LLM pre-training across 1,000+ GPUs. Implement 3D parallelism, manage GPU clusters (SLURM/Kubernetes), optimize InfiniBand/RDMA networking and memory, and automate checkpointing and failure recovery for long training runs.
We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The ideal candidate will have a deep understanding of GPU clusters and extensive experience in system engineering to ensure efficient and reliable training processes.
Responsibilities:
- Orchestrate distributed training runs across 1,000+ GPUs using PyTorch, DeepSpeed, or Megatron-LM.
- Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors.
- Automate checkpointing and failure recovery during month-long training runs.
Required Skills:
- Deep expertise in 3D parallelism (Data, Tensor, Pipeline).
- Experience managing SLURM or Kubernetes-based GPU clusters.
- Strong systems engineering background (C++, CUDA, Python).
Similar Jobs
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Leads the vision, strategy, roadmap, and lifecycle of NBCUniversal’s mobile product experience. Partners with design, engineering, production, analytics, and business teams to define priorities, requirements, metrics, launches, and improvements. Focuses on onboarding, engagement, retention, monetization, experimentation, and long-term product health while communicating tradeoffs and building the mobile product organization.
Top Skills:
ConsoleCross-PlatformMobilePc
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Leads production strategy and delivery planning for an AAA game product pillar. Owns milestone and gate commitments, coordinates pod plans and cross-pillar dependencies, tracks risks and delivery confidence, represents the pillar in executive reviews, evaluates scope changes, and improves planning and delivery practices. Partners with product, design, craft, and production leadership to ensure integrated, achievable outcomes.
Top Skills:
Ai TechnologiesJIRA
Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Design, build, and maintain customer-facing React frontend systems for search, catalog navigation, data lineage, dashboards, and AI-assisted analysis. Improve performance across large trees, tables, and streaming interfaces; shape product and UX decisions with design and product partners; raise engineering quality through reviews, documentation, complex bug resolution, and mentorship.
Top Skills:
Atlassian Design SystemCypressJavaScriptJestPlaywrightReactTypescript
What you need to know about the Toronto Tech Scene
Although home to some of the biggest names in tech, including Google, Microsoft and Amazon, Toronto has established itself as one of the largest startup ecosystems in the world. And with over 2,000 startups — more than 30 percent of the country's total startups — Toronto continues to attract new businesses. Be it helping entrepreneurs manage their finances, simplifying business operations by automating payroll or assisting pharmaceutical companies in launching new drugs, the city's tech scene is just getting started.

.png)
