Genesis Logo

Genesis

Training / AI Infrastructure

Posted 11 Days Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in CA
Senior level
In-Office or Remote
Hiring Remotely in CA
Senior level
Design, build, and optimize distributed PyTorch training systems for multi-node GPU clusters. Profile and eliminate performance bottlenecks across data pipelines to GPU kernels, implement low-level CUDA/cuDNN/Triton kernels, tune CPU/GPU/memory/network utilization, and develop monitoring and debugging tools for large-scale training runs.
The summary above was generated by AI
What You’ll Do
  • Drive down wall-clock time to convergence by profiling and eliminating bottlenecks across the foundation model training stack stack, from data pipelines to GPU kernels

  • Design, build, and optimize distributed training systems (PyTorch) for multi-node GPU clusters, ensuring scalability, robustness, and high utilization

  • Implement efficient low-level code (CUDA, cuDNN, Triton, custom kernels) and integrate it seamlessly into high-level training frameworks

  • Optimize workloads for hardware efficiency: CPU/GPU compute balance, memory management, data throughput, and networking

  • Develop monitoring and debugging tools for large-scale runs, enabling rapid diagnosis of performance regressions and failures

What You’ll Bring
  • Deep experience in distributed systems, ML infrastructure, or high-performance computing (8+ years)

  • Production-grade expertise in Python

  • Low-level performance mastery: CUDA/cuDNN/Triton, CPU–GPU interactions, data movement, and kernel optimization

  • Scaling at the frontier: experience with PyTorch and training jobs using data, context, pipeline, and model parallelism

  • System-level mindset with a track record of tuning hardware–software interactions for maximum utilization

Similar Jobs

Senior level
Agency • Artificial Intelligence • Blockchain • Web3
Design, orchestrate, and optimize large-scale LLM pre-training across 1,000+ GPUs. Implement 3D parallelism, manage GPU clusters (SLURM/Kubernetes), optimize InfiniBand/RDMA networking and memory, and automate checkpointing and failure recovery for long training runs.
Top Skills: 3D ParallelismC++CudaDeepspeedGpuInfinibandKubernetesMegatron-LmPythonPyTorchRdmaSlurm
11 Hours Ago
Remote or Hybrid
Canada
Senior level
Senior level
Cloud • Insurance • Payments • Software • Business Intelligence • App development • Big Data Analytics
The Enterprise Account Executive will focus on acquiring new business in the insurance sector, managing client relationships, and collaborating with sales teams to enhance reach.
Top Skills: CRMSalesforce
12 Hours Ago
Remote or Hybrid
ON, CAN
Senior level
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Lead and execute incident response engagements, perform host and network forensics across Windows, macOS, and Linux, conduct basic malware analysis and reverse engineering, develop hunting methods, produce reports and remediation plans, engage with legal and executives, and contribute thought leadership and public-facing content.
Top Skills: AIAWSAzureBroGCPLinuxmacOSSuricataWindowsZeek

What you need to know about the Toronto Tech Scene

Although home to some of the biggest names in tech, including Google, Microsoft and Amazon, Toronto has established itself as one of the largest startup ecosystems in the world. And with over 2,000 startups — more than 30 percent of the country's total startups — Toronto continues to attract new businesses. Be it helping entrepreneurs manage their finances, simplifying business operations by automating payroll or assisting pharmaceutical companies in launching new drugs, the city's tech scene is just getting started.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account