ELEKS Logo

ELEKS

Infrastructure/GPU Cluster/Platform Operations Lead

Sorry, this job was removed at 03:55 a.m. (PST) on Friday, Sep 11, 2026
Be an Early Applicant
Remote or Hybrid
Hiring Remotely in Canada
Senior level
Remote or Hybrid
Hiring Remotely in Canada
Senior level

Similar Jobs

17 Minutes Ago
Easy Apply
Remote
CAN
Easy Apply
Senior level
Senior level
Artificial Intelligence • Edtech • Machine Learning • Software
Execute and optimize global paid acquisition and remarketing campaigns across Apple Search Ads, Google, Meta, TikTok, and evolving channels. Manage budgets, bidding, targeting, pacing, creative rotation, tracking, and structured testing. Analyze CAC, ROAS, LTV, attribution, and incrementality to improve acquisition and retention. Partner with Creative, Product, Data, and Engineering on campaign performance, landing pages, measurement, and recommendations across international markets.
Top Skills: Apple Search AdsGa4Google AdsLookerMeta AdsTiktok Ads
An Hour Ago
Remote or Hybrid
CA
Senior level
Senior level
eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Develop and manage Square’s always-on B2B content engine across social, thought leadership, customer stories, partner content, and industry narratives. Build content calendars, formats, briefs, campaigns, and AI-assisted workflows; coordinate cross-functional teams and external partners; oversee production through launch; and analyze performance to optimize messaging, formats, and distribution.
Top Skills: Ai ToolsLinkedInSocial AnalyticsSocial Media Platforms
An Hour Ago
Easy Apply
Remote or Hybrid
CA
Easy Apply
Senior level
Senior level
Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Build and operate MITRE ATT&CK-aligned threat detections across cloud, endpoint, identity, email, and application telemetry. Own the detection lifecycle, advance detection-as-code with Git, testing, CI/CD, and observability, and improve data ingestion and enrichment. Evaluate AI-assisted security capabilities, translate threat intelligence into detections, and partner with Security Operations during investigations and incident exercises. The role requires strong Python and SQL skills, SIEM expertise, AWS and Linux knowledge, and independent project leadership.
Top Skills: Agentic FrameworksAirflowAWSChatopsCi/CdContainersDatabricksEdrFedrampGitGoLinuxLlmsMcpMitre Att&CkPythonSIEMSQLTerraform
Lead design and operations of GPU infrastructure and Kubernetes-based AI platform environments. Manage capacity planning, optimize training/inference workloads, define reliability and security practices, implement IaC, monitoring, and disaster recovery, and collaborate with AI engineering to evolve infrastructure.
The summary above was generated by AI
ELEKS is looking for an Infrastructure/GPU Cluster/Platform Operations Lead in Canada.
Alberta-based candidates are strongly preferred (Calgary or Edmonton). Canada-based candidates will also be considered.
 
ABOUT CLIENT

Our customer is building a next-generation AI platform that enables organizations to securely develop, govern, and operationalize artificial intelligence while ensuring that sensitive data and organizational knowledge remain fully under their control. The platform combines advanced AI capabilities with enterprise-grade governance, security, and data sovereignty to support mission-critical decision-making.

The solution serves government organizations and enterprise customers operating in highly regulated and security-sensitive environments, where reliability, accountability, and trust are essential. The platform supports intelligent decision-making across strategic planning, workforce intelligence, and organizational operations, helping customers leverage AI without compromising security, compliance, or control over their data.

REQUIREMENTS

  • 8+ years of Infrastructure Engineering or Platform Operations experience
  • Experience managing GPU clusters for AI workloads
  • Strong Kubernetes administration skills
  • Experience with NVIDIA GPU technologies and CUDA ecosystem
  • Experience with cloud infrastructure (Azure, AWS or GCP)
  • Knowledge of storage, networking, and high-performance computing environments
  • Experience implementing Infrastructure as Code (Terraform or similar)
  • Strong operational leadership skills
  • Experience supporting AI platform infrastructure
  • Upper-Intermediate or higher level of English

RESPONSIBILITIES

  • Lead GPU infrastructure design and operations
  • Manage Kubernetes-based AI platform environments
  • Optimize infrastructure for AI training and inference workloads
  • Define operational standards and reliability practices
  • Collaborate with AI engineering teams
  • Implement monitoring, security, and disaster recovery strategies
  • Lead infrastructure capacity planning
  • Support technical roadmap and infrastructure evolution

What you need to know about the Vancouver Tech Scene

Raincouver, Vancity, The Big Smoke — Vancouver is known by many names, and in recent years, it has gained a reputation as a growing hub for both tech and sustainability. Renowned for its natural beauty, the city has become a magnet for professionals eager to create environmental solutions, and with an emphasis on clean technology, renewable energy and environmental innovation, it's attracted companies across various industries, all working toward a shared goal: advancing clean technology.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account