Design, orchestrate, and optimize large-scale LLM pre-training across 1,000+ GPUs. Implement 3D parallelism, manage GPU clusters (SLURM/Kubernetes), optimize InfiniBand/RDMA networking and memory, and automate checkpointing and failure recovery for long training runs.
We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The ideal candidate will have a deep understanding of GPU clusters and extensive experience in system engineering to ensure efficient and reliable training processes.
Responsibilities:
- Orchestrate distributed training runs across 1,000+ GPUs using PyTorch, DeepSpeed, or Megatron-LM.
- Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors.
- Automate checkpointing and failure recovery during month-long training runs.
Required Skills:
- Deep expertise in 3D parallelism (Data, Tensor, Pipeline).
- Experience managing SLURM or Kubernetes-based GPU clusters.
- Strong systems engineering background (C++, CUDA, Python).
Similar Jobs
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Designs, builds, and operates production-grade agentic AI systems and orchestration frameworks. Responsibilities include prompt architecture, tool and API integrations, monitoring, evaluation, cost controls, reliability improvements, and governance documentation. The engineer collaborates with data science, data engineering, and governance teams to ensure reliable, compliant workflows using structured healthcare and pharmaceutical data.
Top Skills:
LanggraphLlm ApisMlopsPydantic Ai
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Leads the vision, strategy, and execution of AI-powered data and analytics capabilities for Medical Affairs. Oversees patient journey analytics, advanced segmentation, stakeholder evaluation, data architecture, governance, and AI automation frameworks. Evaluates external innovations, builds proprietary solutions, manages analytics delivery teams, and advises senior leadership on data investments, innovation roadmaps, compliance, ethics, and competitive advantage.
Top Skills:
Advanced SegmentationAi/MlAutomation FrameworksData ArchitectureData GovernanceData ScienceHcp Network AnalysisHealthcare Data PlatformsPatient Journey Analytics
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Owns the global operational backbone of the InsightXplorer medical insights platform, including access provisioning, licensing, user administration, request intake and triage, hosting coordination, governance, reporting, metrics, workload visibility, and resource oversight. The role manages competing priorities across regions, coordinates recurring team cadences, supports leadership-ready reporting, and oversees contracted resources while collaborating with technical partners and cross-functional stakeholders.
Top Skills:
Access GovernanceAnalytics PlatformsData Platform OperationsInsightxplorer
What you need to know about the Toronto Tech Scene
Although home to some of the biggest names in tech, including Google, Microsoft and Amazon, Toronto has established itself as one of the largest startup ecosystems in the world. And with over 2,000 startups — more than 30 percent of the country's total startups — Toronto continues to attract new businesses. Be it helping entrepreneurs manage their finances, simplifying business operations by automating payroll or assisting pharmaceutical companies in launching new drugs, the city's tech scene is just getting started.

