Cerebras Systems Inc.

Canada
774 Total Employees
Year Founded: 2015

Jobs at Cerebras Systems Inc.

Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.

Artificial Intelligence • Hardware • Software • Semiconductor
Design, build, and operate CI/CD, Kubernetes-based platforms, deployment automation, and observability for engineering workflows. Improve reliability, performance, and scalability across cloud and on-prem environments, debug cross-boundary failures, perform root-cause analysis, and deliver durable platform software and self-service tooling.
Artificial Intelligence • Hardware • Software • Semiconductor
Operate and scale production AI inference infrastructure, run releases and capacity changes, build self-service CD pipelines and automation, extend telemetry and observability, collaborate on SLOs, post-mortems, and capacity planning to reduce operational toil.
8 Hours AgoSaved
Remote
2 Locations
Artificial Intelligence • Hardware • Software • Semiconductor
Execute hardware bring-up, validation, and telemetry monitoring for Cerebras AI clusters in data centers. Perform power-on sequencing, first-line troubleshooting, log collection, incident support under senior guidance, and contribute feedback to tooling and documentation while learning system architecture and networking fundamentals.
8 Hours AgoSaved
Remote or Hybrid
2 Locations
Artificial Intelligence • Hardware • Software • Semiconductor
Design, implement, and maintain Python frameworks and services that orchestrate distributed engineering workflows across machines and clusters. Build scheduling, execution, resource management, failure recovery, and test infrastructure. Define APIs and abstractions, reason about concurrency and distributed-systems behavior, debug complex multi-system issues, write automated tests and documentation, and partner with platform, CI, release, QA, and product teams to deliver scalable infrastructure.
8 Hours AgoSaved
Remote
2 Locations
Artificial Intelligence • Hardware • Software • Semiconductor
Design and implement system-level debugging, validation, and observability platforms. Build automated anomaly collection/analysis, visualization and root-cause tools, failure classification and monitoring frameworks. Extend compilers, runtimes and programming interfaces for profiling and instrumentation, improve bring-up and low-level debug workflows, lead cross-functional initiatives, support incident response, and establish debuggability and reliability best practices.
8 Hours AgoSaved
Remote
2 Locations
Artificial Intelligence • Hardware • Software • Semiconductor
Design, implement, optimize, and validate high-performance ML and linear algebra kernels for Cerebras hardware. Develop low-level assembly and CSL routines, use mathematical performance models, create unit/system tests, and collaborate with chip and system architects to maximize compute utilization and scale kernels for state-of-the-art AI/HPC workloads.
8 Hours AgoSaved
Remote
2 Locations
Artificial Intelligence • Hardware • Software • Semiconductor
Design and implement high-performance distributed runtime components for large-scale training and inference. Optimize data and communication pipelines, enable scalable multi-node execution, collaborate with ML and compiler teams, diagnose performance issues via profiling, and contribute to system architecture and roadmap for cutting-edge AI workloads.
Artificial Intelligence • Hardware • Software • Semiconductor
Integrate, validate, and productionize cross-stack inference features across AI frameworks, runtime, compiler, kernels, distributed systems, and hardware. Drive zero-to-one projects, debug system-wide failures, manage accelerated timelines, and improve automation, diagnostics, and repeatable integration practices while collaborating across software and hardware teams.
Artificial Intelligence • Hardware • Software • Semiconductor
Lead a hands-on engineering team to improve kernel-centric reliability of large AI compute clusters. Own technical vision, build diagnostic and debug tooling, collaborate with SW and HW teams to reduce downtime, speed failure analysis, and mentor engineers to deliver scalable, reliable production systems.
8 Hours AgoSaved
Remote or Hybrid
2 Locations
Artificial Intelligence • Hardware • Software • Semiconductor
Implement, optimize, and validate high-performance machine learning and linear algebra kernels for the Cerebras Wafer-Scale Engine. Develop low-level kernel routines using the Cerebras Software Language and C++, apply parallel programming techniques, profile and tune performance, debug correctness and hardware utilization issues, build tests and validation, and collaborate with compiler, performance, and hardware engineers.
Artificial Intelligence • Hardware • Software • Semiconductor
Early-career SDET on Release Integration Testing for Cerebras AI Inference Core. Write automation and tests across the AI stack, triage cross-stack failures, maintain branch stability, collect qualification evidence, and collaborate with feature and infra teams to prepare features for production release.
Artificial Intelligence • Hardware • Software • Semiconductor
Implement and scale LLM training, fine-tuning, and post-training techniques (RL-based). Build evaluation and data pipelines, debug ML stack issues, optimize training/inference workflows, and ship maintainable ML infrastructure code.
8 Hours AgoSaved
Remote
2 Locations
Artificial Intelligence • Hardware • Software • Semiconductor
Design and develop software to automate bare-metal configuration, orchestration, scheduling, and job placement for large wafer-scale clusters. Build upgrade/patch workflows, monitoring, HA failure handling, visualization and alerting, and user/admin tools. Ensure on-premise and cloud deployment support and develop tests and debugging tools for distributed systems.
8 Hours AgoSaved
Remote
2 Locations
Artificial Intelligence • Hardware • Software • Semiconductor
Drive end-to-end ML model inference performance: build kernel- and system-level performance models, optimize kernel microcode and compiler algorithms, debug runtime performance on system and cluster, and develop tooling to visualize and analyze performance data from the Wafer Scale Engine and compute cluster.
8 Hours AgoSaved
Remote
2 Locations
Artificial Intelligence • Hardware • Software • Semiconductor
Own quality and reliability of the Cerebras Inference Platform by building test infrastructure and automation, validating Kubernetes and hardware deployments, debugging distributed systems and networking, developing testbeds for performance and scalability, and partnering with platform engineers to ensure production-ready releases.
8 Hours AgoSaved
Remote
2 Locations
Artificial Intelligence • Hardware • Software • Semiconductor
Build, productionize, and optimize a GPU-based inference stack combining GPU prefill with Cerebras decode. Implement and operate model-serving APIs, vLLM/PyTorch/ROCm runtimes, deployment and reliability practices, performance profiling and optimization, cross-layer debugging, numerical validation, and benchmarking/infrastructure for production inference at scale.