Staff Data Engineer
Sonatype
- Canada, Canada
- Remote
- Posted Aug 18, 2026
Job description
About the role
The Staff Data Engineer will join the Data Platform team to design, build and scale data pipelines, models and storage that power analytics, machine learning and business intelligence across Sonatype. The role works closely with product, engineering and business stakeholders, owns parts of the platform on Databricks and Spark, drives architectural vision, mentors engineers and ensures data is reliable, accessible and actionable.
About the company
Sonatype is a software supply chain security company that provides an end‑to‑end solution combining proactive protection against malicious open source, enterprise‑grade SBOM management and open source dependency management. As the founders of Nexus Repository and stewards of Maven Central, it serves over 2,000 organizations—including 70% of the Fortune 100 and 15 million developers—helping them build secure, high‑quality software at scale.
Requirements
- 8+ years of experience as a Data Engineer or similar backend engineering role
- Bachelor’s degree in Computer Science, Engineering, or related technical field
- Experience optimizing Spark jobs, joins, and managing Delta Lake architecture for batch and streaming data
- Experience leveraging AI‑assisted development tools and AI/ML technologies to improve data engineering workflows
- Strong programming skills in Python, Scala, or Java
- Hands‑on experience with distributed data systems such as Spark or Kafka
- Proficiency in writing complex SQL and NoSQL queries and optimizing them for performance
- Experience building and maintaining robust ETL/ELT pipelines in production
- Understanding of data modeling techniques (star schema, dimensional modeling, etc.)
- Familiarity with software supply chain, cybersecurity, or large‑scale software ecosystem data
- Track record of improving data platform reliability, scalability, performance, and cost efficiency
- Experience with workflow orchestration tools like Airflow, Dagster, or similar
- Hands‑on experience with cloud data platforms, particularly AWS
- Familiarity with modern table formats such as Delta Lake, Apache Iceberg, or Apache Hudi
- Experience implementing data observability, lineage, governance, and automated data quality frameworks
- Experience designing real‑time or streaming data architectures using data lake technologies