Cox Exponential Jobs

Founding Engineer, AI Infra

Cox Exponential

Founding Engineer, AI Infra

Reposted 22 Days Ago

Remote or Hybrid

Hiring Remotely in CA

Senior level

Remote or Hybrid

Hiring Remotely in CA

Senior level

Design, build, and operate end-to-end training and inference infrastructure for large language and multimodal models. Improve efficiency (memory, parallelism, kernel optimizations), ensure robust scalable training and RL pipelines, optimize low-latency/high-throughput serving (quantization, caching, speculative decoding), manage multi-GPU and multi-cloud orchestration, and productionize new algorithms with strong observability and reproducibility.

The summary above was generated by AI

About Goaly

At Goaly, our mission is to make custom AI affordable for every business. Our founding team comes from the front lines of top AI labs and tech giants (Meta MSL, TikTok AI, Google DeepMind, xAI, Microsoft Research, etc.), where we built large-scale training infrastructure powering trillion-parameter models and scaled GenAI models to a global user base. Now, we are building something we wish we had before: a platform that makes training and adapting custom AI affordable for all modern companies, not just Big Tech. Our north star is ambitious: for a domain-specific task, reach 90% of SOTA performance at less than 10% of the cost. To get a taste of what we are doing, see our first tech blog.

About the Role

You will sit at the intersection of systems engineering and applied ML, building specialized infrastructure that keeps large language and multimodal models fast, reliable, and cost-effective. You will partner with research, product, and infra teams to ship production-ready platforms for training and serving AI at scale.

Key Responsibilities

Efficiency & performance: Improve LLM training and inference efficiency through better memory utilization, optimized parallelism, and kernel-level innovations (e.g. FlashAttention, CUDA/Triton).
Training & RL robustness: Build scalable, stable training and RL pipelines with strong reproducibility, observability, and debuggability.
Serving & inference optimization: Design and tune high-throughput, low-latency model serving systems, including quantization, caching, and speculative decoding.
Scalability & infrastructure: Own end-to-end training and inference infrastructure — from data ingestion and checkpointing to multi-GPU and multi-cloud orchestration.
Production enablement: Work closely with researchers and product engineers to turn new algorithms into reliable, production-ready systems.

Requirements

5+ years building or operating ML infrastructure at scale, ideally supporting large language or multimodal models.
Deep understanding of GPU architecture, distributed training frameworks (PyTorch, DeepSpeed, Megatron, Ray), and parallelism strategies.
Hands-on experience running inference stacks (vLLM / SGLang, TGI, Triton) and optimizing them via low-level profiling.
Strong software engineering fundamentals in Python and one of C++/Rust/Go, with clean, reliable code shipped to production.
Working knowledge of modern data pipelines, feature stores, and vector databases used in production AI systems.
Comfort automating infrastructure with Kubernetes, Terraform/Pulumi, and observability stacks (Prometheus, Grafana, OpenTelemetry).

Bonus Points

Experience deploying open-source LLMs (Llama 3, Qwen, DeepSeek) or training custom foundation models.
Contributions to ML systems tooling (compilers, kernels, inference runtimes) or open-source infrastructure projects.
Background in reinforcement learning, evaluation harnesses, or alignment tooling that hardens production AI systems.

Similar Jobs

Tapestry - Coach and Kate Spade

Manager, Field Customer Experience

7 Hours Ago

Remote or Hybrid

Toronto, ON, CAN

Senior level

eCommerce • Fashion • Retail • Sales • Wearables • Design

Lead regional retail training and customer experience initiatives, deliver and implement sales, service, and clienteling programs, monitor KPIs, coach store teams and managers, support onboarding and digital tool adoption (Coach Journey, Client Compass), and drive consistent brand service standards across the market.

Top Skills: Client CompassCoach JourneyExcelMS OfficePowerPointWord

Block

Software Engineer

9 Hours Ago

In-Office or Remote

Mid level

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency

Build and operate ingestion, reconciliation, and reporting systems that reconcile card-network and partner settlement files against internal transactions. Deliver end-to-end features for traceability, accounting journals, tax and regulatory reporting, and compliance tooling while ensuring reliability, scalability, and data privacy.

Top Skills: Ai ToolsAirflowAWSBigQueryCi/CdDelta LakeGoHadoopIso-8583JavaKafkaKubernetesObservabilityPysparkPythonSnowflakeSparkSQLTemporalTerraform

Block

Program Manager

9 Hours Ago

In-Office or Remote

Senior level

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency

Lead development and implementation of AI legal and compliance frameworks. Conduct cross-functional AI reviews, maintain governance docs (model cards, impact assessments), partner with engineering on transparency and fairness requirements, build incident response and monitoring processes, and scale AI legal review workflows.

Top Skills: Ai/MlAlgorithmic Impact AssessmentGenerative AiLlmsModel CardsModel Monitoring

What you need to know about the Toronto Tech Scene

Although home to some of the biggest names in tech, including Google, Microsoft and Amazon, Toronto has established itself as one of the largest startup ecosystems in the world. And with over 2,000 startups — more than 30 percent of the country's total startups — Toronto continues to attract new businesses. Be it helping entrepreneurs manage their finances, simplifying business operations by automating payroll or assisting pharmaceutical companies in launching new drugs, the city's tech scene is just getting started.