Equifax Inc. Logo

Equifax Inc.

Agentic AI Optimization Developer

Posted One Month Ago
Be an Early Applicant
In-Office
Toronto, ON, CAN
Mid level
In-Office
Toronto, ON, CAN
Mid level
Build and maintain golden datasets and automated evaluation pipelines to audit multi-agent LLM behavior. Trace trajectories, tune prompts and function-calling, optimize tool integrations (APIs, RPA), monitor production drift and hallucinations, and implement guardrails and RAG pipelines to ensure reliable autonomous workflows.
The summary above was generated by AI

Synopsis of the role

At Equifax, we are moving past passive AI chat interfaces to build the future of autonomous workflows. We are creating intelligent, self-correcting multi-agent systems that can navigate complex software environments, utilize external tools, and solve open-ended business problems with minimal human intervention. To ensure these systems are safe, reliable, and enterprise-grade, we are seeking an analytical Agentic AI Evaluation & Tuning Engineer

In this role, you will be the guardian of our production AI reliability. You will bridge the gap between raw Large Language Model (LLM) capabilities and flawless autonomous execution. Unlike traditional software testers or prompt engineers, you will focus on the behavior, decision-making logic, tool-use efficiency, and long-term stability of multi-agent architectures. Your mission is to build the automated evaluation frameworks that keep our agents accurate, cost-effective, and hallucination-free. 

What you will do

Golden Dataset Curation & Automated Evaluation

  • Build the "Golden Set": Curate, maintain, and augment high-quality reference datasets (Golden Sets) of documents, user queries, and expected agent trajectories to serve as the ultimate source of truth for testing.

  • Automate Eval Cycles: Design and implement automated, continuous evaluation pipelines to measure agent accuracy, latency, token spend, and fallback reliability before code hits production.

  • Trajectory & Reasoning Auditing: Trace and dissect complex, multi-step agent "thought" processes (e.g., ReAct, Reflection loops) to pinpoint exactly where an agent deviates from its intended logic path.

Agent Tuning & Developer Collaboration

  • Behavioral Optimization: Refine system prompts, context windows, and few-shot examples to optimize how agents execute complex, multi-step workflows.

  • Tool & Function-Calling Optimization: Fine-tune how agents interact with external APIs, databases, and UiPath RPA workflows—minimizing execution errors, redundant calls, and token overhead.

  • Augment Development: Partner closely with AI Solution Leads and AI Agent Developers to feed evaluation insights back into the development lifecycle, helping them build robust, reusable, and self-correcting agent components.

Production Guardrails & Lifecycle Management (LLMOps)

  • Defeat Drift & Hallucinations: Actively monitor deployed agents to identify, troubleshoot, and mitigate semantic drift, prompt injections, infinite execution loops, and hallucinations.

  • Maintain Autonomous Integrity: Implement robust guardrail frameworks to ensure agents maintain reliable, fact-based autonomous decision-making post-deployment in production.

  • RAG & Knowledge Integration: Optimize Domain-Specific Knowledge Bases and Retrieval-Augmented Generation (RAG) pipelines to ensure agents pull from accurate data rather than assumptions.

What Experience You Need

  • Experience: 3+ years of professional experience in software quality engineering, test automation, or data/ML engineering, with a dedicated focus on LLM testing, prompt tuning, or orchestration patterns over the last 1–2 years. 

  • Agentic & LLM Frameworks: Proven hands-on experience working with LLM orchestration frameworks (e.g., LangGraph, ADKs or specialized internal SDKs).

  • Function Calling Mastery: Deep understanding of JSON schema design for LLM tool-calling, function-calling, and structured outputs.

  • Advanced Debugging & Automation: Strong background in writing automated test scripts (Python-heavy) and using tracing/observability concepts to debug cascading errors in asynchronous, non-deterministic systems.

What Could Set You Apart

  • Experience with AI evaluation and observability platforms 

  • Live production experience testing Agentic workflows and GenAI solutions 

  • Familiarity with Google Cloud AI suite (Vertex & Gemini Enterprise Agent Platform) and UiPath ecosystem (Maestro).

  • Experience utilizing LLMs to securely generate high-quality synthetic data for edge-case testing.

  • Proficiency in Python or TypeScript, with a deep understanding of asynchronous programming, API design, and microservices architecture.

  • Demonstrated learning agility and a proactive approach to mastering new technologies. 

 This is a newly created position.

Primary Location:

CAN-Toronto-5700 Yonge

Function:

Function - Tech Dev and Client Services

Schedule:

Full time

Equifax Inc. Toronto, Ontario, CAN Office

Toronto, Canada

Similar Jobs

17 Minutes Ago
Hybrid
Hamilton, ON, CAN
Mid level
Mid level
Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Plans, develops, and executes capital projects for a food manufacturing facility. Responsibilities include developing budgets and cash-flow forecasts, coordinating technical solutions, supporting equipment and process strategies, managing cross-functional stakeholders, ensuring engineering standards and regulatory requirements are met, and delivering projects according to quality, safety, environmental, and business objectives.
Top Skills: Capital Project ManagementCodes And StandardsElectrical ControlsEngineering StandardsLocal Regulations
19 Minutes Ago
Remote or Hybrid
CA
Senior level
Senior level
eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Own Square Hardware sales and revenue across major U.S. retail partners. Build and scale retail partnership strategies, manage existing retailer and distributor relationships, source and negotiate new partnerships, lead cross-functional launches and forecasting, and manage retail marketing budgets through test-and-learn investments.
An Hour Ago
Remote or Hybrid
ON, CAN
Junior
Junior
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Sells CrowdStrike’s cybersecurity platform to commercial customers in Manitoba and Saskatchewan. Owns the full sales cycle from prospecting through closing, generates net-new business, manages quota, forecasts in Clari, and collaborates with sales development, engineering, channel, and marketing teams. The role requires consultative selling to mid-market and enterprise organizations, executive-level engagement, partner strategy, competitive positioning, and use of sales and AI tools. Occasional travel and modified hours may be required.
Top Skills: Ai TechnologiesClariCloudCybersecurity SolutionsGongLinkedin Sales NavigatorSaaSSalesforce (Sfdc)Zoominfo

What you need to know about the Toronto Tech Scene

Although home to some of the biggest names in tech, including Google, Microsoft and Amazon, Toronto has established itself as one of the largest startup ecosystems in the world. And with over 2,000 startups — more than 30 percent of the country's total startups — Toronto continues to attract new businesses. Be it helping entrepreneurs manage their finances, simplifying business operations by automating payroll or assisting pharmaceutical companies in launching new drugs, the city's tech scene is just getting started.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account