Bot Auto Logo

Bot Auto

Senior Software Engineer, Developer Platform

Posted Yesterday
In-Office or Remote
Hiring Remotely in CA
Senior level
In-Office or Remote
Hiring Remotely in CA
Senior level
Design, deploy, and operate large-scale workflow orchestration platforms (Argo/Airflow/hybrid) for simulation, ML training, data pipelines and CI/CD. Build SDKs, abstractions, and self-service tooling; ensure reliability, observability, cost-efficiency, GPU and cross-cluster scheduling on Kubernetes across cloud and on-prem; integrate with storage, streaming, registries, and CI/CD; document best practices and mentor engineers.
The summary above was generated by AI
Company Introduction

At Bot Auto, we are revolutionizing the transportation of goods with our cutting-edge autonomous trucks, enhancing the quality of life for communities around the globe. With the agility of a start-up and the wisdom of seasoned experts, Bot Auto boasts a team that has achieved numerous world-firsts and unparalleled innovations. United by a shared vision, we create miracles and propel the future of transportation. Join us and transform your dreams into reality.

We are seeking a highly skilled and motivated Senior Software Engineer to architect, build, and operate the workflow orchestration platforms that power Bot Auto's engineering and autonomy workloads. From simulation and machine learning training to data pipelines and CI/CD, our teams depend on reliable, scalable workflow systems to move fast. In this role, you will own one or a hybrid of orchestration platforms (such as Argo Workflows and Airflow), operate them at scale, and develop the internal platforms and abstractions built on top of them that make running complex workloads simple, observable, and cost-efficient.

Key Responsibilities
  • Architect, deploy, and operate workflow orchestration platforms (e.g., Argo Workflows, Airflow, or a hybrid) supporting simulation, machine learning and model training, data pipelines, CI/CD, and other general-purpose workloads.
  • Build internal platforms, abstractions, SDKs, and self-service tooling on top of orchestration engines to make authoring, running, and monitoring workflows simple and reliable for engineers.
  • Operate workflow platforms at scale on Kubernetes across cloud (AWS) and on-prem data center environments, handling scheduling, autoscaling, GPU and heterogeneous resources, and cross-cluster orchestration.
  • Ensure reliability, performance, and cost efficiency of workloads through observability, queuing and prioritization, retries, and resource optimization.
  • Partner with ML, simulation, data, and infrastructure teams to understand workload requirements and deliver fit-for-purpose pipelines.
  • Integrate workflow platforms with storage, data streaming and event systems, artifact and model registries, and CI/CD tooling.
  • Establish best practices, templates, and documentation for workflow authoring and operations; mentor engineers across the company.
  • Handle user-impacting issues promptly with clear communication — mitigate in the short term and follow up with durable long-term solutions.
Required Qualifications
  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent experience
  • 5+ years of hands-on experience in platform engineering, infrastructure, DevOps, or SRE roles
  • Significant experience with workflow orchestration platforms such as Argo Workflows, Airflow, or comparable systems
  • Strong software development skills in one or more languages: Python, Go, Java, or JavaScript/TypeScript
  • Solid understanding of Kubernetes and distributed systems
Preferred Qualifications
  • Expert-level experience operating Argo Workflows, Airflow, and/or other engines (e.g., Prefect, Dagster, Temporal, Kubeflow Pipelines, Flyte)
  • Experience orchestrating ML training, simulation, or large-scale data and batch workloads, including GPU scheduling
  • In-depth Kubernetes experience (EKS, GKE, AKS, RKE2/Rancher) and cross-cluster orchestration
  • IaC tools proficiency, including Terraform, Pulumi, OpenTofu, or Ansible
  • Experience with data streaming and event platforms, including NATS JetStream, Kafka, Pulsar, or RabbitMQ
  • Familiarity with observability stacks: Prometheus, Grafana, Loki, OpenTelemetry, or comparable
  • Demonstrated ability to optimize workload cost and performance without compromising reliability

Similar Jobs

5 Days Ago
Easy Apply
Remote
Canada
Easy Apply
Senior level
Senior level
Big Data • Fintech • Mobile • Payments • Financial Services
Lead delivery of ML feature platform initiatives: design, build, and operate backend systems for feature creation, storage, backfilling and serving. Collaborate with product, design, and analytics, own quarterly goals, improve platform reliability and code quality, mentor engineers, and support on-call and operational monitoring.
Top Skills: AWSKotlinKubernetesMySQLPython
19 Days Ago
Easy Apply
Remote
Canada
Easy Apply
Senior level
Senior level
Big Data • Fintech • Mobile • Payments • Financial Services
Build, operate, and extend Affirm's data platform and application tooling to enable self-service data applications at scale. Implement platform features (provisioning, CI/CD, access control), integrations, data governance, automation, observability, and on-call operations. Contribute to roadmap for semantic layers, agentic data tooling, and AI-native workflows while raising engineering and operational standards.
Top Skills: Agentic Data ToolsBigQueryBuildkiteCi/CdDatabricksDockerGithub ActionsLlmsOauthOidcPythonRbacRest ApisSecrets ManagementSemantic LayerSnowflakeSQLWebhooks
14 Days Ago
Remote
Québec, QC, CAN
Senior level
Senior level
Software
Design, build, and maintain scalable backend services for a production ML platform handling millions of LLM calls and predictions daily. Improve observability, reliability, and ease-of-use; collaborate with scientists and engineers to deploy agentic workflows and realtime inference; participate in coding, testing, and deployment.
Top Skills: AWSAws BedrockJavaLitellmLlm Inference GatewaysOpensearch ServerlessPythonSpringTerraformVector Databases

What you need to know about the Toronto Tech Scene

Although home to some of the biggest names in tech, including Google, Microsoft and Amazon, Toronto has established itself as one of the largest startup ecosystems in the world. And with over 2,000 startups — more than 30 percent of the country's total startups — Toronto continues to attract new businesses. Be it helping entrepreneurs manage their finances, simplifying business operations by automating payroll or assisting pharmaceutical companies in launching new drugs, the city's tech scene is just getting started.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account