Cerebras Systems Inc.
Jobs at Cerebras Systems Inc.
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Recently posted jobs
Artificial Intelligence • Hardware • Software • Semiconductor
Operate, monitor, and optimize large-scale AI compute clusters built on the Wafer-Scale Engine. Develop and own software for cluster operations (monitoring, automation, APIs, dashboards), deploy and debug Docker-based services, troubleshoot distributed systems, maximize compute capacity, and participate in a 24/7 on-call rotation while collaborating cross-functionally to improve operational visibility and reliability.
Artificial Intelligence • Hardware • Software • Semiconductor
Own quality and reliability of the Cerebras Inference Platform by building test infrastructure and automation, validating Kubernetes and hardware deployments, debugging distributed systems and networking, developing testbeds for performance and scalability, and partnering with platform engineers to ensure production-ready releases.
Artificial Intelligence • Hardware • Software • Semiconductor
Harden low-level platform layers across firmware, bootloader, Linux kernel, and hardware roots of trust. Design secure/measured boot and attestation, review kernel modules/drivers for memory safety and privilege issues, build kernel telemetry (eBPF), and drive fleet-wide vulnerability response and remediation.
Artificial Intelligence • Hardware • Software • Semiconductor
Lead architecture and evolution of enterprise, data center, and cloud networks for hyperscale AI; design secure, high-performance fabrics (spine/leaf, RDMA); build AI-agent review frameworks; implement segmented zero-trust network designs; set resiliency, observability, and capacity standards; mentor global engineers and guide cross-functional initiatives.
Artificial Intelligence • Hardware • Software • Semiconductor
Design, build, and operate firewall, segmentation, and zero-trust controls across data center, corporate, and cloud networks. Implement infrastructure-as-code for policy and CI/CD deployment, manage lifecycle and rule hygiene, build network detection capabilities, operate VPN/remote access patterns, and document architecture and runbooks while partnering with Network Engineering and Security Operations.
Artificial Intelligence • Hardware • Software • Semiconductor
Join IT & Security to secure and scale enterprise IT, cloud, network, and infrastructure; build automation and tooling; support detection, response, and vulnerability management; improve identity, endpoint, and systems security; and partner cross-functionally to develop processes that enable secure, reliable, and scalable AI workloads.
Artificial Intelligence • Hardware • Software • Semiconductor
Build and maintain customer-facing AI cloud frontends (inference, training, admin consoles, and APIs). Design responsive, high-reliability interfaces to handle high traffic, using modern web frameworks and best practices. Collaborate on architecture, ensure production-grade quality, and thrive in ambiguous, fast-changing environments while sharing knowledge and communicating clearly.
Artificial Intelligence • Hardware • Software • Semiconductor
Design, build, and operate CI/CD, Kubernetes-based platforms, deployment automation, and observability for engineering workflows. Improve reliability, performance, and scalability across cloud and on-prem environments, debug cross-boundary failures, perform root-cause analysis, and deliver durable platform software and self-service tooling.
Artificial Intelligence • Hardware • Software • Semiconductor
Operate and scale production AI inference infrastructure, run releases and capacity changes, build self-service CD pipelines and automation, extend telemetry and observability, collaborate on SLOs, post-mortems, and capacity planning to reduce operational toil.
Artificial Intelligence • Hardware • Software • Semiconductor
Execute hardware bring-up, validation, and telemetry monitoring for Cerebras AI clusters in data centers. Perform power-on sequencing, first-line troubleshooting, log collection, incident support under senior guidance, and contribute feedback to tooling and documentation while learning system architecture and networking fundamentals.
Artificial Intelligence • Hardware • Software • Semiconductor
Design and implement system-level debugging, validation, and observability platforms. Build automated anomaly collection/analysis, visualization and root-cause tools, failure classification and monitoring frameworks. Extend compilers, runtimes and programming interfaces for profiling and instrumentation, improve bring-up and low-level debug workflows, lead cross-functional initiatives, support incident response, and establish debuggability and reliability best practices.
Artificial Intelligence • Hardware • Software • Semiconductor
Design, implement, and maintain Python frameworks and services that orchestrate distributed engineering workflows across machines and clusters. Build scheduling, execution, resource management, failure recovery, and test infrastructure. Define APIs and abstractions, reason about concurrency and distributed-systems behavior, debug complex multi-system issues, write automated tests and documentation, and partner with platform, CI, release, QA, and product teams to deliver scalable infrastructure.
Artificial Intelligence • Hardware • Software • Semiconductor
Design, implement, optimize, and validate high-performance ML and linear algebra kernels for Cerebras hardware. Develop low-level assembly and CSL routines, use mathematical performance models, create unit/system tests, and collaborate with chip and system architects to maximize compute utilization and scale kernels for state-of-the-art AI/HPC workloads.
Artificial Intelligence • Hardware • Software • Semiconductor
Design and implement high-performance distributed runtime components for large-scale training and inference. Optimize data and communication pipelines, enable scalable multi-node execution, collaborate with ML and compiler teams, diagnose performance issues via profiling, and contribute to system architecture and roadmap for cutting-edge AI workloads.
Artificial Intelligence • Hardware • Software • Semiconductor
Integrate, validate, and productionize cross-stack inference features across AI frameworks, runtime, compiler, kernels, distributed systems, and hardware. Drive zero-to-one projects, debug system-wide failures, manage accelerated timelines, and improve automation, diagnostics, and repeatable integration practices while collaborating across software and hardware teams.
Artificial Intelligence • Hardware • Software • Semiconductor
Lead a hands-on engineering team to improve kernel-centric reliability of large AI compute clusters. Own technical vision, build diagnostic and debug tooling, collaborate with SW and HW teams to reduce downtime, speed failure analysis, and mentor engineers to deliver scalable, reliable production systems.
Artificial Intelligence • Hardware • Software • Semiconductor
Implement, optimize, and validate high-performance machine learning and linear algebra kernels for the Cerebras Wafer-Scale Engine. Develop low-level kernel routines using the Cerebras Software Language and C++, apply parallel programming techniques, profile and tune performance, debug correctness and hardware utilization issues, build tests and validation, and collaborate with compiler, performance, and hardware engineers.
Artificial Intelligence • Hardware • Software • Semiconductor
Early-career SDET on Release Integration Testing for Cerebras AI Inference Core. Write automation and tests across the AI stack, triage cross-stack failures, maintain branch stability, collect qualification evidence, and collaborate with feature and infra teams to prepare features for production release.
Artificial Intelligence • Hardware • Software • Semiconductor
Implement and scale LLM training, fine-tuning, and post-training techniques (RL-based). Build evaluation and data pipelines, debug ML stack issues, optimize training/inference workflows, and ship maintainable ML infrastructure code.
Artificial Intelligence • Hardware • Software • Semiconductor
Design and develop software to automate bare-metal configuration, orchestration, scheduling, and job placement for large wafer-scale clusters. Build upgrade/patch workflows, monitoring, HA failure handling, visualization and alerting, and user/admin tools. Ensure on-premise and cloud deployment support and develop tests and debugging tools for distributed systems.
