Sonatype Logo

Sonatype

Staff Data Engineer

Posted One Month Ago
Remote
Hiring Remotely in Canada
Senior level
Remote
Hiring Remotely in Canada
Senior level
Design, build, and maintain scalable data pipelines, optimize Databricks/Spark and Delta Lake architectures, deliver trusted datasets for analytics and ML, implement observability and data quality, and drive platform architecture and team mentorship.
The summary above was generated by AI

Sonatype is the software supply chain security company. We provide the world’s best end-to-end software supply chain security solution, combining the only proactive protection against malicious open source, the only enterprise grade SBOM management and the leading open source dependency management platform. This empowers enterprises to create and maintain secure, quality, and innovative software at scale.

As founders of Nexus Repository and stewards of Maven Central, the world’s largest repository of Java open-source software, we are software pioneers and our open source expertise is unmatched. We empower innovation with an unparalleled commitment to build faster, safer software and harness AI and data intelligence to mitigate risk, maximize efficiencies, and drive powerful software development.

More than 2,000 organizations, including 70% of the Fortune 100 and 15 million software developers, rely on Sonatype to optimize their software supply chains.

About the role:

    We’re looking for a Staff Data Engineer to join our growing Data Platform team. You’ll play a key role in designing and scaling the infrastructure and pipelines that power analytics, machine learning, and business intelligence across Sonatype.You’ll work closely with stakeholders across product, engineering, and business teams to ensure data is reliable, accessible, and actionable. This role is ideal for someone who thrives on solving complex data challenges at scale and enjoys building high-quality, maintainable systems.

    At Sonatype, we:
  • Use data with purpose: you'll get the chance to work on problems that directly impact how the world builds secure software
  • Use modern tooling: you'll get the chance to leverage the best of open-source and cloud-native technologies
  • Have a deep collaborative culture: you'll be joining a passionate team that values learning, autonomy, and impact

What you'll do:

  • Design, build, and maintain scalable data pipelines and ETL/ELT processes

  • Architect and optimize data models and storage solutions for analytics and operational use

  • Collaborate with data scientists, analysts, and engineers to deliver trusted, high-quality datasets

  • Own and evolve parts of our data platform using Databricks and Spark

  • Implement observability, alerting, and data quality monitoring for critical pipelines

  • Drive best practices in data engineering, including documentation, testing, and CI/CD

  • As a Staff Engineer you will help drive long-term architectural vision and mentor the team on engineering best practices, while partnering with stakeholders to ensure data solutions support business outcomes.

  • Contribute to the design and evolution of our next-generation data lakehouse architecture

What you bring:

  • 8+ years of experience as a Data Engineer or in a similar backend engineering role

  • Bachelor’s degree in Computer Science, Engineering, or a related technical field

  • Databricks Optimization: Tune Spark jobs, optimize join performance, and manage Delta Lake architecture for batch and streaming data.

  • Experience leveraging AI-assisted development tools and AI/ML technologies to improve data engineering workflows, developer productivity, data quality and ops.

  • Strong programming skills in Python, Scala, or Java

  • Hands-on experience with distributed data systems like Spark or Kafka

  • Proficient in writing complex SQL and NoSQL queries and optimizing queries for performance

  • Experience building and maintaining robust ETL/ELT pipelines in production

  • Understanding of data modeling techniques (star schema, dimensional modeling, etc.)

It’d be great if you also had:

  • Familiarity with software supply chain, cybersecurity, or large-scale software ecosystem data

  • A track record of improving data platform reliability, scalability, performance, and cost efficiency

  • Familiarity with workflow orchestration tools (Airflow, Dagster, or similar)

  • Hands-on experience with cloud data platforms, particularly AWS

  • Familiarity with modern table formats such as Delta Lake, Apache Iceberg, or Apache Hudi

  • Experience implementing data observability, lineage, governance, and automated data quality frameworks

  • Experience designing real-time or streaming data architectures using data lake technologies

Things we are proud of:

  • 2026 Gartner® Magic Quadrant™ Leader for Software Supply Chain Security
  • 2026 Celebrating 15 Years of Sonatype Research Labs – Industry-leading software supply chain and open source security research
  • 2026 Founding Member of the Linux Foundation Initiative for Open Source Sustainability
  • 2026 State of the Software Supply Chain® Report – Continuing industry leadership in software supply chain security and AI security research
  • 2025 Visionary in Gartner® Magic Quadrant™ for Application Security Testing!
  • 2025 AI Compliance Solution of the Year - AI Breakthrough Awards
  • 2025 DEVIES Award to our SBOM Manager for a new product for its innovation and impact in developer technology
  • 2024 Industry Leader in Forrester-Wave for Software Composition Analysis (2024 Q4 report)
  • Constellation AST Shortlist: Sonatype has been listed on the Constellation ShortList™ for Application Security Testing for 2024
  • Data Breakthrough Awards: Sonatype was announced as a 2024 winner in the "Open Source Data Solution of the Year."
  • SD Times: Best in Show Security
  • Fast Company Best Workplaces for Innovators 2024
  • The Herd Top 100 Private Software Companies 2024
  • Diversity & Inclusion Working Groups
  • Parental Leave Policy
  • Paid Volunteer Time Off (VTO)

At Sonatype, we value diversity and inclusivity. We offer perks such as parental leave, diversity and inclusion working groups, and flexible working practices to allow our employees to show up as their whole selves. We are an equal-opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. If you have a disability or special need that requires accommodation, please do not hesitate to let us know.
 

Similar Jobs

4 Days Ago
Remote
Canada
Senior level
Senior level
Professional Services • Consulting
Own and evolve data infrastructure, distributed pipelines, data modeling, analytics, governance, semantic layers, and customer-facing data products. Operate cloud data systems, support production reliability and on-call incident response, and develop streaming and knowledge-management capabilities. The role also evaluates platform tools, manages stakeholder needs, and contributes to platform design and zero-to-one data system development.
Top Skills: BigQueryCloud InfrastructureDatabricksDbtInfrastructure As CodeSnowflakeStreaming Data Systems
7 Days Ago
In-Office or Remote
2 Locations
Expert/Leader
Expert/Leader
Software
Leads the architecture, development, reliability, governance, and scalability of batch and real-time data platforms. Designs pipelines, self-service tooling, observability frameworks, and disaster recovery practices; mentors data engineers; establishes engineering standards; and partners cross-functionally to enable trusted data access for analytics, data science, and product teams. The role also contributes to strategic roadmaps, platform innovation, technical debt reduction, and on-call support.
Top Skills: APIsAWSChange Data Capture (Cdc)Ci/CdCloudFormationContainerizationData SerializationDimensional ModelingGraphQLInfrastructure As CodeLoggingMonitoringObservabilityPyTorchRayReactRuby On RailsSnowflakeTensorFlowTerraformTracingTypescriptWorkflow Orchestration
8 Days Ago
Remote
Canada
Expert/Leader
Expert/Leader
Software • Energy • Utilities
Lead the architecture and scaling of Overstory’s data platform for AI-powered vegetation analysis. Design and operate large-scale geospatial and temporal data pipelines, orchestration systems, event-driven workflows, service-oriented architecture, data quality frameworks, and observability practices. Drive architectural standards, system design, scalability, reliability, and traceability while mentoring engineers and collaborating with data, ML, and product teams.
Top Skills: AirflowBigQueryDagsterGCPPrefectPub/SubPython

What you need to know about the Toronto Tech Scene

Although home to some of the biggest names in tech, including Google, Microsoft and Amazon, Toronto has established itself as one of the largest startup ecosystems in the world. And with over 2,000 startups — more than 30 percent of the country's total startups — Toronto continues to attract new businesses. Be it helping entrepreneurs manage their finances, simplifying business operations by automating payroll or assisting pharmaceutical companies in launching new drugs, the city's tech scene is just getting started.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account