Huawei Canada Logo

Huawei Canada

Researcher - Reinforcement Learning

Reposted 12 Days Ago
Be an Early Applicant
In-Office
Markham, ON, CAN
Expert/Leader
In-Office
Markham, ON, CAN
Expert/Leader
The role involves advancing reinforcement learning techniques for LLMs, including designing training pipelines and evaluating agentic behaviors, contributing to scientific publications.
The summary above was generated by AI

Huawei Canada has an immediate 12-month contract opening for a Reinforcement Learning Researcher.


About the team:

Founded in 2012, the Noah’s Ark lab has evolved into a prominent research organization with notable achievements in academia and industry. The lab’s mission focuses on advancing artificial intelligence and related fields to benefit the company and society. Driven by impactful, long-term projects, the aim is to enhance state-of-the-art research while integrating innovations into the company's products and services, including LLMs, RL, NLP, computer vision, AI theory, and Autonomous driving.

About the job:

  • Enabling Large Language Models (LLMs) to learn from experience, interaction, and environment feedback, moving beyond static fine-tuning toward continual, agentic self-improvement.

  • LLM post-training paradigms (e.g., RLHF, GRPO, reward-free methods, etc.).

  • Agentic reinforcement learning for tool-using and browsing-based LLMs trained in interactive environments.

  • Agentic evaluation and benchmarking, including design of multi-turn, verifiable reasoning tasks.

  • Your work will involve implementing and evaluating new training and evaluation pipelines for reasoning-enhanced LLMs and tool-using agents, scaling experiments on large GPU clusters, and contributing to scientific insights and publications in this emerging area.

About the ideal candidate:

  • PhD degree in Computer Science or related fields or master's degree with comparable experience.

  • Strong foundation in deep learning, including architectures such as Transformers and optimization techniques for large models.

  • Practical or research experience in reinforcement learning, self-supervised learning, or language model fine-tuning.

  • Proven research record in AI by having at least one paper as the first author in top tier venues, such as NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, ICRA.

  • Solid proficiency in Python and experience with PyTorch, DeepSpeed, Megatron and other distributed training frameworks.

  • Familiarity with LLM post-training pipelines (RLHF, GRPO/PPO, SFT, LoRA, MoE, etc.) is an asset.

  • Experience with multi-agent RL, tool-use / browser/coding agents, is an asset.

  • Strong communication and writing skills; enthusiasm for open research and collaborative problem-solving.

Huawei aims to support a French-speaking work environment for its employees in Quebec. We have taken steps to avoid requiring a language other than French for this position. However, proficiency in English is essential for this role for the following reasons:

The person will be required to communicate regularly with colleagues located outside Quebec, where English is the primary language used for communication between offices. In addition, the nature of the tasks related to this position, which falls within a highly specialized field of artificial intelligence, also requires knowledge of English.

HQ

Huawei Canada Markham, Ontario, CAN Office

19 Allstate Pky, Markham, Ontario, Canada, L3R 5A4

Similar Jobs

An Hour Ago
Hybrid
Mid level
Mid level
Automotive • Hardware • Robotics • Software • Transportation • Manufacturing
Provide day-to-day HR support to managers and employees including full-cycle recruitment, onboarding, employee relations, performance management, attendance and disability administration, HRIS (Dayforce) recordkeeping, benefits/payroll support, compliance with employment legislation, and partnering on workforce planning, training, and organizational development.
Top Skills: DayforceHrisExcelMS OfficeMicrosoft PowerpointMicrosoft WordWorkday
An Hour Ago
Hybrid
Senior level
Senior level
Automotive • Hardware • Robotics • Software • Transportation • Manufacturing
Lead cradle-to-grave automotive programs: plan and control schedules, budgets, suppliers and cross-functional teams; drive APQP, launch readiness, customer negotiations, divisional metrics, and continuous improvement while ensuring safety and quality compliance.
Top Skills: ApqpCmsMqsPdpSorSow
An Hour Ago
Hybrid
Newmarket, ON, CAN
Internship
Internship
Automotive • Hardware • Robotics • Software • Transportation • Manufacturing
Support design, build, test, and validation of electromechanical automotive components and subsystems. Tasks include prototype assembly, test-fixture and vehicle integration, bench work (soldering, wiring), using LabVIEW and measurement tools, basic CAE/hand calculations, documentation (DVP&R, BOM, GD&T), and assisting with analyses (FEA, FMEA) under senior engineer guidance.
Top Skills: ArduinoCC++CadCanDfm/DfaDoeFeaFmeaGd&TI2CLabviewLinMultimetersOscilloscopesPcbPythonSpiUart

What you need to know about the Toronto Tech Scene

Although home to some of the biggest names in tech, including Google, Microsoft and Amazon, Toronto has established itself as one of the largest startup ecosystems in the world. And with over 2,000 startups — more than 30 percent of the country's total startups — Toronto continues to attract new businesses. Be it helping entrepreneurs manage their finances, simplifying business operations by automating payroll or assisting pharmaceutical companies in launching new drugs, the city's tech scene is just getting started.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account