Cerebras Systems Inc.
Jobs at Cerebras Systems Inc.
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Recently posted jobs
Artificial Intelligence • Hardware • Software • Semiconductor
Build and evolve production ML inference APIs and serving infrastructure across GPU and Cerebras accelerator backends. Responsibilities include API design, model integration, request routing, disaggregated prefill/decode coordination, performance optimization, correctness validation, observability, testing, SDKs, and developer tooling. The role requires close collaboration across compiler, runtime, cloud, product, and customer-facing teams to deliver reliable, compatible, scalable inference capabilities.
Artificial Intelligence • Hardware • Software • Semiconductor
Designs deployment-ready AI data center infrastructure, including rack layouts, equipment placement, bills of materials, power allocations, port maps, and cabling. Adapts cluster designs to site constraints, validates integrated compute, storage, networking, and power systems, supports deployment readiness, resolves integration issues, and automates design and validation workflows using scripts and reusable tools.
Artificial Intelligence • Hardware • Software • Semiconductor
Own and scale major areas of Cerebras’s Developer Console, including billing, usage tracking, quotas, request logs, and metrics. Design frontend systems and backend services using Next.js, TypeScript, GraphQL, Postgres, and Redis. Make architectural decisions across APIs, data models, distributed systems, and real-time processing. Lead projects from conception through production, improve reliability and observability, partner with product and design, and mentor engineers.
Artificial Intelligence • Hardware • Software • Semiconductor
Build and operate distributed infrastructure software for Cerebras clusters. Responsibilities include bare-metal automation, Kubernetes operators, workload scheduling, gRPC control-plane services, authorization, observability pipelines, failure detection, high availability, automated recovery, and fleet-facing CLIs and APIs. The role requires production Go and Python, deep Kubernetes expertise, strong Linux and networking debugging skills, and experience operating distributed systems at scale.
Artificial Intelligence • Hardware • Software • Semiconductor
Bring up and validate next-generation AI hardware systems and supporting software. Debug hardware-software integration failures using logs, telemetry, and diagnostic tools; develop automation frameworks, testing infrastructure, and observability tooling; reproduce and triage complex issues; and collaborate with hardware engineers to qualify systems for production.
Artificial Intelligence • Hardware • Software • Semiconductor
Designs and operates secure, scalable cloud infrastructure and identity platforms across AWS and data centers. Builds IAM, authentication, authorization, SSO, MFA, provisioning, secrets, and key management solutions. Develops Terraform-based automation using Python or Go, supports Kubernetes services, implements Zero Trust controls, and improves reliability, observability, and developer experience. The role includes architecture reviews, cross-team technical leadership, mentoring, production support, incident response, and on-call participation.
Artificial Intelligence • Hardware • Software • Semiconductor
Designs and leads CI/CD, artifact lifecycle, build, test, release, and developer infrastructure systems. Improves cloud infrastructure, code review workflows, branching strategies, automation, observability, and engineering productivity. Troubleshoots complex distributed systems, drives architectural improvements, participates in on-call and incident response, and mentors engineers across the Developer Productivity organization. Contributes to AI tooling that automates repetitive engineering workflows.
Artificial Intelligence • Hardware • Software • Semiconductor
Design and deliver production FPGA solutions for chassis-to-wafer IO including RoCE v2/RDMA network interfaces, switching fabric, and serial IO. Develop RTL, produce bitstreams, optimize bandwidth and latency, debug large AI cluster networking, coordinate board bringup, and lead cross-functional projects with DV, embedded software, and architecture teams.
Artificial Intelligence • Hardware • Software • Semiconductor
Operate, monitor, and optimize large-scale AI compute clusters built on the Wafer-Scale Engine. Develop and own software for cluster operations (monitoring, automation, APIs, dashboards), deploy and debug Docker-based services, troubleshoot distributed systems, maximize compute capacity, and participate in a 24/7 on-call rotation while collaborating cross-functionally to improve operational visibility and reliability.
Artificial Intelligence • Hardware • Software • Semiconductor
Build and maintain the inference orchestration layer across datacenter clusters. Design and implement highly available, low-latency systems, define platform direction (Kubernetes CRDs/operators), lead production incident response, optimize performance and capacity, drive observability and security, and collaborate with ML, product, and cloud teams to productionize scalable inference services.
One Month AgoSaved
Artificial Intelligence • Hardware • Software • Semiconductor
Design, build, and operate CI/CD, Kubernetes-based platforms, deployment automation, and observability for engineering workflows. Improve reliability, performance, and scalability across cloud and on-prem environments, debug cross-boundary failures, perform root-cause analysis, and deliver durable platform software and self-service tooling.
Artificial Intelligence • Hardware • Software • Semiconductor
Lead testing strategy and execution for ML API features, building scalable test frameworks and automation to validate accuracy, fairness, and performance of inference systems. Drive pre-deployment and production validation across distributed cloud, multi-region, and hardware-backed inference environments. Mentor SDETs, champion best practices, debug complex distributed issues, and lead cross-functional quality initiatives to improve test coverage and engineering efficiency.
Artificial Intelligence • Hardware • Software • Semiconductor
New graduate software engineer working on low-level, performance-oriented software for AI accelerator systems. Design, implement, test, and debug hardware-near components, support system bring-up, performance optimization, tooling, and cross-functional collaboration with hardware, firmware, compiler, and infrastructure teams.
Artificial Intelligence • Hardware • Software • Semiconductor
Bring up state-of-the-art open-source or proprietary LLMs on Cerebras CSX systems across the full stack: model translation, graph lowering, compiler optimizations, runtime integration, performance tuning, debugging, and prototyping tooling improvements to accelerate future bring-ups.
Artificial Intelligence • Hardware • Software • Semiconductor
Build and maintain CI/CD pipelines and artifact lifecycle systems, provision and optimize cloud (AWS) CI infrastructure, troubleshoot build and pipeline issues, improve developer tooling and build/test infrastructure, support code review/workflow improvements, participate in on-call incident response, and contribute to AI tooling to boost engineering productivity.
Artificial Intelligence • Hardware • Software • Semiconductor
Design, build, and operate secure, scalable cloud infrastructure and identity platforms across AWS and data centers. Implement IAM/IGA, SSO/MFA, federation protocols, secrets and key management, and Zero Trust controls. Automate with Terraform, Python, and Go, deploy on Kubernetes, participate in on-call and incident response, and collaborate with Security and Engineering to deliver secure-by-design systems.
Artificial Intelligence • Hardware • Software • Semiconductor
Build and productionize a GPU-accelerated prefill serving path for LLM inference: deploy, operate, and optimize model-serving, runtime (vLLM/PyTorch/ROCm), multi-node GPU infrastructure, reliability, performance, correctness, and benchmarking.
Artificial Intelligence • Hardware • Software • Semiconductor
Develop and maintain infrastructure to build, test, operate, simulate, and evaluate Cerebras software. Design scalable automation workflows for cloud and datacenter, manage CI/CD, containerized deployments, and monitor software quality and performance on the Wafer Scale Engine.
Artificial Intelligence • Hardware • Software • Semiconductor
Lead capacity planning and fleet strategy for the Inference Service: build rolling forecasts, support datacenter bring-up, decide model/cluster allocation, run utilization reporting, drive capacity tool adoption, coordinate cross-functional execution, and manage incident postmortems.
Artificial Intelligence • Hardware • Software • Semiconductor
Lead automation and platform engineering to eliminate toil and deliver self-service GitOps-driven CD, capacity provisioning, and observability for large-scale inference clusters. Define SLOs/SLIs, mentor SREs, support incident escalation, implement reliability practices, and measure impact via deployment velocity, SLO compliance, MTTR, and adoption of self-service workflows.
