Robots & Pencils Logo

Robots & Pencils

Staff ML Engineer – AWS Trainium & SageMaker

Posted 13 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in Canada
Expert/Leader
Remote
Hiring Remotely in Canada
Expert/Leader
Train and operate production machine learning models on Amazon SageMaker using AWS Trainium. Optimize PyTorch code for Trainium compilation, memory, throughput, and distributed training; diagnose hardware, compiler, device, data, and code issues; build cost-aware end-to-end training pipelines; and collaborate directly with client and internal engineering teams to deliver production workloads.
The summary above was generated by AI
Robots & Pencils is an AWS Partner building production AI systems for enterprise clients who need real engineering, not proofs of concept that never ship. We work forward-deployed, embedded directly with client teams, solving the problems that are too new or too specialized for a typical vendor relationship to handle.

The Role
We're looking for an engineer who can operate and train models on Amazon SageMaker running on AWS Trainium, AWS's custom silicon built specifically for large-scale model training. This isn't a role where you call an API and wait. You'll be walking up the stack: understanding what a training request actually looks like at the Trainium hardware and compiler level, then carrying that understanding all the way up through PyTorch training code and into a production SageMaker pipeline.
PyTorch is the backbone of this work. If you know the framework deeply and you're comfortable reasoning about how your code actually behaves on custom accelerator hardware rather than treating it as a black box, this role is built around that skill set specifically.
 
What You'll Do
  • Train and operate models on Amazon SageMaker with AWS Trainium as the underlying compute
  • Write and optimize PyTorch training code with a real understanding of how it compiles and executes on Trainium (NeuronCore architecture, compiler behavior, memory and throughput tradeoffs)
  • Diagnose training run issues that show up specifically because of the hardware, not just the model, distinguishing a data or code problem from a compiler or device-level one
  • Translate a request for "a Trainium job" into an actual working, cost-aware training pipeline, end to end
  • Tune distributed training runs for throughput and cost on SageMaker's training infrastructure
  • Work directly with client and internal engineering teams to scope and deliver real production training workloads, not experiments that stay in a notebook
What You'll Bring
  • Strong, hands-on PyTorch experience, ideally including distributed or multi-device training
  • Production experience with Amazon SageMaker for training and/or inference
  • Comfort working close to the hardware layer: you understand device-specific compilation and can debug issues that are actually about the accelerator, not just the model
  • AWS Trainium or Inferentia (Neuron SDK) experience is a strong plus; if you don't have it yet but have deep PyTorch and a track record of picking up new hardware targets fast, we want to talk to you
  • Solid Python fundamentals and comfort operating in a client-facing, production engineering environment

Similar Jobs

5 Hours Ago
Remote
Canada
Mid level
Mid level
Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Build and operate end-to-end full-stack features for Loom billing, subscriptions, integrations, and platform migrations. Develop responsive React interfaces, scalable backend APIs and microservices, payment and identity integrations, and reliable cloud-based systems. Responsibilities include observability, monitoring, testing, code reviews, documentation, security, performance optimization, and participation in Agile ceremonies and technical design reviews.
Top Skills: AWSBitbucketCi/CdConfluenceCypressDockerGitJavaJavaScriptJestJIRAMicroservicesMochaMySQLNode.jsPostgresPythonReactRest ApisTypescript
Yesterday
Easy Apply
Remote or Hybrid
Canada
Easy Apply
Junior
Junior
Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Serve as the primary post-implementation contact for top customers, craft joint success plans, run executive business reviews and workshops, mentor teammates, support French-speaking customers, and advise on customizing Samsara’s IoT platform to drive safety, efficiency, and sustainability.
Top Skills: IotSamsara PlatformVehicle TelematicsVideo-Based Safety
Yesterday
Easy Apply
Remote or Hybrid
Canada
Easy Apply
Expert/Leader
Expert/Leader
Marketing Tech • Social Media • Software • Analytics • Business Intelligence
Own the strategy, roadmap, delivery, pricing, and growth of Sprout Social’s Listening product. Drive adoption and competitive differentiation through customer discovery, market analysis, product investments, and cross-functional collaboration with Engineering, Design, Marketing, Sales, Customer Experience, and GTM teams. Establish success metrics, optimize usage and pricing opportunities, communicate product direction, and lead roadmap execution across distributed teams.
Top Skills: AnalyticsProduct AnalyticsSaaSSocial Media Intelligence

What you need to know about the Vancouver Tech Scene

Raincouver, Vancity, The Big Smoke — Vancouver is known by many names, and in recent years, it has gained a reputation as a growing hub for both tech and sustainability. Renowned for its natural beauty, the city has become a magnet for professionals eager to create environmental solutions, and with an emphasis on clean technology, renewable energy and environmental innovation, it's attracted companies across various industries, all working toward a shared goal: advancing clean technology.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account