Veeda AI Logo

Veeda AI

Senior AI Infrastructure Engineer — HPC & Compute Clusters

Posted Yesterday
Remote or Hybrid
Hiring Remotely in Zürich
Senior level
Remote or Hybrid
Hiring Remotely in Zürich
Senior level
Design, deploy, and maintain large-scale bare-metal GPU clusters for distributed training. Tune Slurm schedulers and Kubernetes integrations, optimize storage and high-speed networking, and build observability pipelines to ensure reliable, high-throughput GPU workloads for distributed deep learning.
The summary above was generated by AI
About US

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.

Responsibilities
  • GPU Cluster Orchestration: Design, deploy, and maintain large-scale bare-metal GPU clusters running Slurm and Kubernetes to power multi-node distributed training runs.

  • Workload Management & Scheduler Tuning: Configure, tune, and optimize Slurm schedulers, dynamic job queuing, priority topologies, and autoscaling for optimal GPU utilization and job throughput.

  • Kubernetes Integration: Manage K8s clusters and hybrid Slurm-K8s environments (e.g., KubeRay, MPI Operator, Slurm on K8s) for serving, data processing pipelines, and interactive research environments.

  • Storage & Network Performance: Optimize high-speed interconnects (InfiniBand/RoCE, NCCL, NVLink) and high-throughput distributed storage (Lustre, WEKA, Ceph) to keep tens of thousands of GPUs continuously saturated.

  • Reliability & Monitoring: Build observability pipelines (Prometheus, Grafana, DCGM) to proactively detect hardware degradation, GPU silent errors, network flaps, and node failures before they impact training jobs.

Requirements
  • You have a Bachelor's degree in Computer Science, Computer Engineering, or equivalent hands-on experience in high-performance computing (HPC) or infrastructure engineering.

  • You have deep hands-on expertise administering Linux-based HPC clusters running Slurm or Kubernetes (job scheduling, cgroups, fair-share policies, topology configuration).

  • You possess strong troubleshooting skills in low-level Linux networking, kernel parameters, hardware diagnostics, and storage systems (e.g., NFS, NVMe-oF, or distributed filesystems).

  • You are proficient in automation and infrastructure-as-code tools (e.g., Ansible, Terraform, Helm, Python/Bash scripting).

  • Background supporting distributed deep learning frameworks (DeepSpeed, Ray).

Nice to Have
  • Experience managing high-density GPU infrastructure (NVIDIA H100/B200/B300 systems, DGX/HGX architectures, DCGM, NCCL tuning).

  • Experience with high-speed network fabrics including InfiniBand (Subnet Manager, SMDB) and RoCE (v2).

Similar Jobs

4 Hours Ago
In-Office or Remote
Senior level
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Lead HEOR and evidence strategy for obesity and internal medicine early-pipeline assets. Design and oversee RWD studies, economic models, and patient-reported outcomes; influence trial design; manage evidence generation, publications, vendor relationships, and cross-functional stakeholder engagement to secure reimbursement and patient access.
Top Skills: ChatgptMicrosoft Copilot
4 Hours Ago
In-Office or Remote
Senior level
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Lead global HEOR and evidence-generation strategy for obesity assets. Design and execute RWE, economic models, and PRO strategies; create launch dossiers; manage cross-functional teams, vendors, budgets, and external partnerships to support reimbursement and patient access.
Top Skills: ChatgptMicrosoft Copilot
10 Hours Ago
Remote
Senior level
Senior level
Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Cybersecurity • Data Privacy
Own technical relationship for Swiss enterprise and mid-market accounts: run discovery, demos, and technical validations (POC), architect Rubrik solutions across on‑prem, cloud, SaaS, identity and AI data workflows, qualify opportunities, present to technical and CxO stakeholders, and partner with account teams to close deals.
Top Skills: Ai-Driven Data WorkflowsBackup And Disaster RecoveryCloudCyber ResilienceData ProtectionData ResilienceIdentity SecurityObservabilityRemediationRubrikRubrik Agent CloudRubrik Security CloudSaaS

What you need to know about the Vancouver Tech Scene

Raincouver, Vancity, The Big Smoke — Vancouver is known by many names, and in recent years, it has gained a reputation as a growing hub for both tech and sustainability. Renowned for its natural beauty, the city has become a magnet for professionals eager to create environmental solutions, and with an emphasis on clean technology, renewable energy and environmental innovation, it's attracted companies across various industries, all working toward a shared goal: advancing clean technology.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account