Genesis Logo

Genesis

Training / AI Infrastructure

Posted 2 Days Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in CA
Senior level
In-Office or Remote
Hiring Remotely in CA
Senior level
Design, build, and optimize distributed PyTorch training systems for multi-node GPU clusters. Profile and eliminate performance bottlenecks across data pipelines to GPU kernels, implement low-level CUDA/cuDNN/Triton kernels, tune CPU/GPU/memory/network utilization, and develop monitoring and debugging tools for large-scale training runs.
The summary above was generated by AI
What You’ll Do
  • Drive down wall-clock time to convergence by profiling and eliminating bottlenecks across the foundation model training stack stack, from data pipelines to GPU kernels

  • Design, build, and optimize distributed training systems (PyTorch) for multi-node GPU clusters, ensuring scalability, robustness, and high utilization

  • Implement efficient low-level code (CUDA, cuDNN, Triton, custom kernels) and integrate it seamlessly into high-level training frameworks

  • Optimize workloads for hardware efficiency: CPU/GPU compute balance, memory management, data throughput, and networking

  • Develop monitoring and debugging tools for large-scale runs, enabling rapid diagnosis of performance regressions and failures

What You’ll Bring
  • Deep experience in distributed systems, ML infrastructure, or high-performance computing (8+ years)

  • Production-grade expertise in Python

  • Low-level performance mastery: CUDA/cuDNN/Triton, CPU–GPU interactions, data movement, and kernel optimization

  • Scaling at the frontier: experience with PyTorch and training jobs using data, context, pipeline, and model parallelism

  • System-level mindset with a track record of tuning hardware–software interactions for maximum utilization

Similar Jobs

Senior level
Agency • Artificial Intelligence • Blockchain • Web3
Design, orchestrate, and optimize large-scale LLM pre-training across 1,000+ GPUs. Implement 3D parallelism, manage GPU clusters (SLURM/Kubernetes), optimize InfiniBand/RDMA networking and memory, and automate checkpointing and failure recovery for long training runs.
Top Skills: 3D ParallelismC++CudaDeepspeedGpuInfinibandKubernetesMegatron-LmPythonPyTorchRdmaSlurm
3 Hours Ago
Remote
Canada
Expert/Leader
Expert/Leader
Artificial Intelligence • Software
Lead full-cycle strategic sales in Canada for Fieldguide, building relationships with major audit and advisory firms. Drive net-new revenue, orchestrate multi-regional account strategy, manage complex technical sales cycles, deliver executive presentations, coordinate cross-functional stakeholders, develop international playbooks, and represent the company at industry events. Expected travel up to 30%.
3 Hours Ago
Remote or Hybrid
CA
Senior level
Senior level
Gaming
Lead development of core client systems and engine-level infrastructure in Unity/C#. Diagnose and fix cross-cutting performance issues, upgrade third-party SDKs, evolve build tooling, produce design docs, scope multi-week infrastructure initiatives, mentor engineers, and drive profiling, optimization, and architecture changes to enable future features.
Top Skills: AndroidAWSC#DatadogGitiOSKubernetesUnity

What you need to know about the Vancouver Tech Scene

Raincouver, Vancity, The Big Smoke — Vancouver is known by many names, and in recent years, it has gained a reputation as a growing hub for both tech and sustainability. Renowned for its natural beauty, the city has become a magnet for professionals eager to create environmental solutions, and with an emphasis on clean technology, renewable energy and environmental innovation, it's attracted companies across various industries, all working toward a shared goal: advancing clean technology.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account