Tali AI Logo

Tali AI

Senior/Staff Applied AI Engineer

Posted 11 Days Ago
Be an Early Applicant
Hybrid
Toronto, ON, CAN
Senior level
Hybrid
Toronto, ON, CAN
Senior level
Owns production AI systems end to end, including evaluation pipelines, agents, retrieval, routing, model selection, data feedback loops, and observability. Diagnoses failures across audio, prompts, models, and infrastructure; improves quality through experiments and production debugging; and leads safe rollouts. The role establishes applied AI engineering standards, translates clinical problems into measurable plans, and delivers reliable systems used by healthcare clinicians.
The summary above was generated by AI
About Tali

Clinicians are drowning in paperwork. We're giving them their time back.

Tali AI is one of the fastest-growing startups in Canada, on a mission to make healthcare more accessible with AI. We're building the clinical operating system: an ambient AI scribe, a billing agent, scheduling, clinical decision support, and more on one platform that allows clinicians to look at their patients instead of their keyboards. Thousands of clinicians across Canada and the US already use Tali, across dozens of specialties, integrated deeply with the fragmented landscape of North American health-record systems.

The traction is real: multiple commercial product lines, 15M+ patient visits documented on the platform, 70+ years worth of time saved for clinicians in under three years. We move fast in a market that rewards speed, in a domain where breaking clinician trust is not recoverable.

The role

You own AI systems end to end. The models, the prompts, the retrieval, the data, the evals, and the debugging when a clinician says the output was wrong.

The ambient scribe listens to a visit and writes the clinical note. A recommender suggests the billing codes a clinician can claim for that visit. Medical search answers a clinical question from real sources. And a growing number of agents run inside the company and in the product. You work across all of it.

Evaluation carries the same weight here as any product feature. It is what makes the rest of it trustworthy.

What you'll work on
  • Build the evaluation pipelines that decide whether the AI is good enough to ship. Automated judges, regression suites, human review, and the datasets underneath them. On the audio path that means word and speaker error rates and the audio-quality measures that tell you a recording was worth trusting.
  • Diagnose failure modes and fix them. The lever might be the prompt, the retrieval, the routing, the model, the audio capture, or a fine-tune.
  • Build production agents. Tool use, orchestration, guardrails, and recovery when a step fails. You also build what sits under them: search, vector storage, and the harness the agents run in.
  • Debug one visit end to end, then find every case like it. You trace a single interaction from audio to delivered note and work out what broke. Then you slice the warehouse to size the problem and prove the fix.
  • Decide which model serves which request, and change that safely. Weighted routing, staged rollout, and attribution good enough that you know which change moved the number.
  • Own these systems in production. You get the alert when quality slips, you find the cause, and you decide what ships to fix it.
  • Turn a vague clinical complaint into a problem statement, a metric, and a plan the team can act on.
  • Set the bar for how Tali does applied AI. Your evals become the evals everyone else runs.
What we're looking for
  • 5+ years in production ML, applied AI, or research engineering. You have owned something that ran for real users and stayed up.
  • Deep evaluation experience. You have built graders, regression suites, or judge pipelines, and you know how to tell when a judge is fooling you.
  • Agentic systems. Multiple models, tool calls, and retrieval, with the failure recovery that makes them safe to run.
  • Strong systems engineering. Backend services, data pipelines, and enough observability that you can answer questions about production quickly.
  • Data-centric instincts. You improve an AI system by improving its data and its feedback loops, and you can say when a prompt change is the smaller lever.
  • Python, plus modern ML tooling. You write code others can run.
  • Candour. You give hard feedback on a colleague's design, and you take it on your own without going quiet.
  • You make the case for the harder right answer in engineering terms and in business terms, then you ship it and own the result.
  • You raise the people around you. Your review makes the next engineer's system better.

This is a senior or staff role depending on your track record. The levelling conversation happens at the end of the interview process.

Bonus points

  • Speech recognition or real-time audio.
  • A regulated domain, such as healthcare or finance.
  • Clinical experience of any kind.

Rigor and shipped systems matter more to us than the domain you learned them in.

What you'll work with
  • Python and TypeScript.
  • GCP and Cloud Run.
  • Vertex AI, Claude, and other frontier LLM and ASR providers.
Is this you?

You're passionate about the application of technology in our users' lives. You default to the simplest way to deliver value to users. And you know when the simple answer stops being good enough. Then you make the technical case and the business case for the harder one, and you ship it. Both halves matter. A technical case with no business case is an unfunded idea. A business case with no technical case is a guess dressed up as a pitch.

If you don't have ground truth data, you mine production for weak labels and generate what you can. When you hit a real ceiling, you cost out annotators, write the case for the investment, build the quality checks, and run the program.

If audio quality is a deep issue for some users, you squeeze what you can from the signal, and you know when and how to make the case for shipping microphones to users. Then you drive the execution.

You build our eval tooling even when it's unglamorous. You debug traces, you talk to clinicians about what went wrong, and you know when to invest in better infrastructure to slash toil.

What success looks like
  • 3 months: You own a real AI problem and you're driving it with minimal oversight. Anyone can now see how good that system is, and how good it was last week.
  • 6 months: Evaluation and rollout are a system you helped build. Experiments repeat, regressions get caught before a clinician sees them, and the team ships faster because of it.
  • 1 year: Tali's AI systems have taken a step change in capability. The next generation of agents is in production. Clinicians trust the output because the rigor behind it is real and repeatable. When a model or eval problem gets hard, people come to you.
Working at Tali

Benefits

Flexible work hours
Comprehensive health and wellness coverage from day one, including unmetered wellness days
Competitive PTO, including winter shutdown Dec 25 - Jan 1, birthdays and Taliversaries, and 'extra long' long weekends
$2000 annually in "Knowledge Dollars" to learn, grow, and level up
Quarterly socials & company outings that bring our team together beyond the day-to-day

Our Core Values

⚡️ Bold: we embrace ambitious goals, make courageous decisions, and take calculated risks to drive impactful innovation and growth
🧩 Resourceful: we're self-directed problem solvers; navigating obstacles, learning and acquiring new skills and making sound judgement calls. We consistently deliver on commitments while maintaining a high standard of quality and dependability
📣 Candid: Being, honest, transparent, and open in all interactions, fostering a culture of trust and authenticity
💛 Caring: Actively supporting and empathizing with our people - customers, patients, and colleagues to help them thrive and achieve their goals
Tali Online
🔗 Tali is one of Linkedin's Top Startups of 2025
🔗 Tali is part of the renowned Digital Supercluster Project
🔗 Check out Tali's CEO, Mahshid Yassaei on Cherry Health's Leaders in Healthcare Podcast
HQ

Tali AI Toronto, Ontario, CAN Office

Hanson St, Toronto, Ontario, Canada

Similar Jobs

39 Minutes Ago
Hybrid
Expert/Leader
Expert/Leader
Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
Leads the technical strategy, roadmap, architecture, and global delivery of cloud-native Core Network User Plane sub-services. Oversees engineering practices including CI/CD, SRE, DevOps, security, compliance, scalability, reliability, and cost efficiency across incubation, trials, commercial deployment, and scale-up. Builds engineering culture, establishes ways of working and performance metrics, and collaborates with global leadership, architects, engineers, and SRE teams to deliver customer-centric telecommunications products.
Top Skills: 3Gpp4G/5G Core NetworksAgileAWSAzureCi/CdCloud-Native ArchitectureDevOpsDistributed CloudDpdkEbpfEdge ComputingGCPHyperscalersPublic CloudSite Reliability Engineering (Sre)Sr-IovTraffic SteeringUpf
8 Hours Ago
Hybrid
Mississauga, ON, CAN
Entry level
Entry level
eCommerce • Fashion • Retail • Sales • Wearables • Design
Leads store operations for Kate Spade New York, supporting sales performance, customer experience, team development, and achievement of business goals. The role requires strong interpersonal skills, strategic agility, creativity, customer focus, and the ability to provide feedback and build effective teams. Candidates must regularly lift 25 pounds, occasionally lift up to 50 pounds, and maneuver throughout the sales floor and stockroom.
Yesterday
Hybrid
Toronto, ON, CAN
Entry level
Entry level
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Lead Data Scientist analyzing global fraud data to identify patterns, develop predictive models, and prevent payment fraud. The role builds fraud-related products and services, extracts insights from large datasets, collaborates with stakeholders, and presents findings to management and clients. Required skills include Python, Spark, complex data queries, and advanced data analysis. Experience with agentic AI systems and Unix is preferred, along with payments or fraud expertise.
Top Skills: SparkLlm-Based AgentsPythonUnix

What you need to know about the Toronto Tech Scene

Although home to some of the biggest names in tech, including Google, Microsoft and Amazon, Toronto has established itself as one of the largest startup ecosystems in the world. And with over 2,000 startups — more than 30 percent of the country's total startups — Toronto continues to attract new businesses. Be it helping entrepreneurs manage their finances, simplifying business operations by automating payroll or assisting pharmaceutical companies in launching new drugs, the city's tech scene is just getting started.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account