Cox Exponential Jobs

Research Scientist/Engineer, Efficient ML Systems

Cox Exponential

Research Scientist/Engineer, Efficient ML Systems

Reposted 24 Days Ago

Remote or Hybrid

Hiring Remotely in CA

Mid level

Remote or Hybrid

Hiring Remotely in CA

Mid level

Research and build efficient ML systems for large-scale LLMs and agentic RL: design algorithms and system techniques, prototype in training/inference stacks, run large-scale experiments, and translate findings into production or publications.

The summary above was generated by AI

About Goaly

At Goaly, our mission is to make custom AI affordable for every business. Our founding team comes from the front lines of top AI labs and tech giants (Meta MSL, TikTok AI, Google DeepMind, xAI, Microsoft Research, etc.), where we built large-scale training infrastructure powering trillion-parameter models and scaled GenAI models to a global user base. Now, we are building something we wish we had before: a platform that makes training and adapting custom AI affordable for all modern companies, not just Big Tech. Our north star is ambitious: for a domain-specific task, reach 90% of SOTA performance at less than 10% of the cost. To get a taste of what we are doing, see our first tech blog.

About the Role

As an AI Research Scientist (Efficient ML Systems) at Goaly, you will research and build the systems that make frontier-scale models practical. This role sits at the intersection of algorithms, systems, and hardware efficiency.You will design and evaluate new training and inference techniques, prototype them in real systems, and push them to production-scale workloads.

Your work will either ship directly into our core platform or lead to publications at top venues such as NeurIPS, ICML, ICLR, or CVPR.This is not a paper-only role. You will write real systems code, run large-scale experiments, and directly shape how modern LLMs and RL systems are trained and deployed.

Core Responsibilities

Research efficient ML systems: Invent and evaluate algorithms and system techniques that improve LLM and agentic RL training and inference efficiency (memory, compute, communication, and stability).
Scale agentic RL: Design and optimize large-scale agentic RL pipelines, including asynchronous training, experience management, reward modeling, and long-horizon stability.
End-to-end experimentation: Design large-scale experiments spanning model architecture, training algorithms, distributed systems, and hardware-aware optimization.
System-aware research: Prototype research ideas directly in training and inference stacks (e.g., parallelism strategies, attention kernels, RL training pipelines) and validate them at scale.
Production & publication: Translate successful ideas into production-ready systems and/or publish them at top-tier conferences with full internal support.

Requirements

Ph.D. or Master's degree in CS, AI, Systems, or related fields (Exceptional undergraduates with strong research capabilities may be considered).
Strong foundation in LLM or large-scale ML training, including Transformers, attention mechanisms, distributed training, and optimization methods.
Experience or strong interest in agentic RL or large-scale reinforcement learning systems, including stability, scalability, or long-horizon training challenges.
Demonstrated interest in efficiency-focused research, such as training acceleration, memory optimization, parallelism, kernels, or RL system robustness.
Proficient in PyTorch or JAX. Clean coding style and strong command of Python.
Adaptability: A fast learner with a strong sense of responsibility, capable of wearing multiple hats and handling cross-stack challenges.

Bonus Points

First-author publications at top conferences (NeurIPS, ICML, ICLR, CVPR, ACL).
High-star open-source projects on Hugging Face or Gold/Silver medals in Kaggle competitions.

Similar Jobs

CrowdStrike

Analyst I, Falcon Complete (Remote, PST/MST)

4 Hours Ago

Remote or Hybrid

Entry level

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity

Detect, contain, and remediate security incidents across Windows, macOS, and Linux. Perform malware and forensic analysis, develop detection and remediation processes, produce customer-facing reports and recommendations, and contribute to public thought leadership. Use scripting/programming and AI tools to enhance investigations and response.

Top Skills: .NetAi TechnologiesCC#Forensic Analysis ToolsLinuxmacOSMalware AnalysisNetwork Analysis ToolsPerlPythonRuby On RailsVbWindows

Coinbase

Senior Software Engineer

6 Hours Ago

Easy Apply

Remote

Canada

Easy Apply

Senior level

Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3

Lead design and delivery of reliability projects for Coinbase's platform: secure service configuration and secrets management, improve canary-based deployments, partner with core services to increase scalability and reduce incidents, drive reliability best practices, and participate in on-call rotations.

Top Skills: AWSAzureDatadogGCPGoKibanaRubyTerraform

Block

Strategic Programs Lead

6 Hours Ago

In-Office or Remote

Senior level

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency

Lead end-to-end delivery of high-priority revenue programs, coordinating Product, Sales, Marketing, Finance, Legal, Risk, and Operations. Drive launch readiness for products and partnerships, establish governance and operating cadences, create repeatable launch frameworks and executive communications, and continuously improve cross-functional execution.

What you need to know about the Vancouver Tech Scene

Raincouver, Vancity, The Big Smoke — Vancouver is known by many names, and in recent years, it has gained a reputation as a growing hub for both tech and sustainability. Renowned for its natural beauty, the city has become a magnet for professionals eager to create environmental solutions, and with an emphasis on clean technology, renewable energy and environmental innovation, it's attracted companies across various industries, all working toward a shared goal: advancing clean technology.