Design, orchestrate, and optimize large-scale LLM pre-training across 1,000+ GPUs. Implement 3D parallelism, manage GPU clusters (SLURM/Kubernetes), optimize InfiniBand/RDMA networking and memory, and automate checkpointing and failure recovery for long training runs.
We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The ideal candidate will have a deep understanding of GPU clusters and extensive experience in system engineering to ensure efficient and reliable training processes.
Responsibilities:
- Orchestrate distributed training runs across 1,000+ GPUs using PyTorch, DeepSpeed, or Megatron-LM.
- Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors.
- Automate checkpointing and failure recovery during month-long training runs.
Required Skills:
- Deep expertise in 3D parallelism (Data, Tensor, Pipeline).
- Experience managing SLURM or Kubernetes-based GPU clusters.
- Strong systems engineering background (C++, CUDA, Python).
Similar Jobs
Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Big Data Analytics • Automation
No responsibilities were provided in the job text. Based on the title, the role likely leads a Support Center of Excellence, defines support strategy and processes, drives performance and best practices, partners with stakeholders, and manages a support organization.
Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Own named enterprise accounts/territory: develop strategic sales plans, build C-level relationships, negotiate complex contracts, collaborate cross-functionally, forecast, close deals, and drive expansion across products.
Top Skills:
CRM
Fintech • Financial Services
Lead technical architecture and roadmap for analytics engineering: design semantic/metrics layer, build AI-ready Snowflake data foundation, optimize dbt projects, mentor engineers, partner with Data Science/Engineering, and champion data governance, observability, and production readiness for AI and downstream analytics.
Top Skills:
Agent FrameworksCi/CdDbtFivetranKafkaObservability ToolsPythonRagSnowflakeSnowflake CortexSnowflake IntelligenceSQLStreaming/Cdc
What you need to know about the Vancouver Tech Scene
Raincouver, Vancity, The Big Smoke — Vancouver is known by many names, and in recent years, it has gained a reputation as a growing hub for both tech and sustainability. Renowned for its natural beauty, the city has become a magnet for professionals eager to create environmental solutions, and with an emphasis on clean technology, renewable energy and environmental innovation, it's attracted companies across various industries, all working toward a shared goal: advancing clean technology.



