Job Summary
About the Role
Our client, a fast-growing data intelligence company serving a major retail/manufacturing industry, is looking for a Senior Data Engineer to own data quality at scale as the company grows. The company processes 600GB+ of SKU-level data daily across a multi-trillion-dollar industry, with production systems operating at terabyte-scale volume. This role sits at the center of ensuring the accuracy, consistency, and reliability of that data, for example correctly identifying and reconciling product variants (e.g., ensuring a “blue pen” isn’t inconsistently mislabeled as a “whiteboard pen” across sources). This person will be hands-on with the existing pipelines and infrastructure that move and validate data at this scale, while also working directly with clients and internal stakeholders to understand and resolve data issues.
This is a high-visibility role for someone who wants ownership over both the technical quality of the data and the relationships built around it. We’re specifically looking for someone who genuinely thrives in ongoing data quality and data governance work long-term. This is sustained, high-scale data quality ownership, not a short-term cleanup project.
Key Responsibilities
– Maintain and optimize existing data pipelines that support large-scale, high-velocity data ingestion (600GB+ processed daily, terabyte-scale production systems)
– Develop and refine dimensional data models using dbt and ClickHouse
– Own data quality processes: detect, diagnose, and resolve inconsistencies (duplicate/mislabeled entities, schema drift, missing or conflicting values) through automated testing and observability
– Build proprietary metrics and KPIs in collaboration with sales and customer success
– Implement statistical methods for anomaly detection (Prophet, ARIMA, Tukey’s Fences)
– Create internal Hex applications and Preset visualizations to democratize data access
– Automate core business processes for efficiency gains
– Enhance modeling infrastructure for scalability and performance
– Partner directly with clients and internal teams to understand data issues, explain root causes, and manage resolution timelines
– Present insights and recommendations to senior leadership
– Stay ahead of industry trends and introduce relevant tools and techniques
Qualifications & Skills
– Bachelor’s degree in a technical field (Economics, Information Systems, or related) preferred
– Strong experience with dbt, SQL, and data warehousing (Snowflake, BigQuery, ClickHouse preferred)
– Proficiency in Python, Docker, and Kubernetes
– Hands-on experience with ETL tools, Spark, and CI/CD pipelines (Git, Jenkins)
– Experience with data visualization tools like Tableau or Hex
– Experience with cloud infrastructure (AWS, Azure, or GCP) and infrastructure-as-code tools (e.g., Terraform)
– Familiarity with data streaming/event pipelines (e.g., Kafka, Kinesis, or similar)
– Comfort working with entity resolution / data matching / deduplication logic
– Strong analytical, problem-solving, and communication skills
– Ability to work cross-functionally in a fast-paced environment
Soft Skills & Traits
– Strong communicator who’s comfortable being client-facing, able to explain technical data issues to non-technical stakeholders clearly and patiently
– Project management instincts: can prioritize competing data issues, set realistic timelines, and keep stakeholders updated without hand-holding
– High tolerance for detail-oriented, sometimes repetitive quality work, genuinely energized by “keeping the data clean” rather than looking to move off it quickly
– Prior experience working at large data scale (terabyte-level or high-volume production data) strongly preferred; this role is demanding specifically because of the scale and velocity involved
– Self-directed and comfortable owning a problem space end to end in a fast-growing, fully remote environment
– Collaborative, works well across engineering, customer success, and leadership
Nice to Have
– Experience in retail, CPG, or e-commerce data environments
– Background supporting AI/ML-adjacent products or pipelines
Share This Job
ABOUT US
Our High Trail Verified members include elite Technical and Solutions Architects, Developers, Engineers, Administrators, Business Analysts, Project Managers and Leaders, with a proven track record of delivering solutions to complex problems.
If you are ready to Elevate Your Work, you’ve come to the right place. Contact us to schedule a consultation.
Other Jobs you might like…
Herndon, Virginia
$170,000 - $195,000
Overview Our client is seeking a Senior Manager – Oracle Fusion Cloud ERP / Solution Architect to serve as the…
McLean, Virginia
$70 - $85
Job Summary Securities Operations Research Analytics & Development (SORAD) – Alteryx Developer / Process Automation Engineer Position Overview We are…
Remote
$60 - $75
Salesforce Product Owner Position Overview We are seeking a Salesforce Product Owner to support a growing Salesforce team by partnering…
Remote
$112,000 - $146,000
Our client, a reputable financial institution is seeking an experienced AVP of Digital Marketing with a focus on Salesforce Marketing…
Remote
$0 - $110,000
A leading organization in the asset management software sector is seeking an experienced Salesforce Business Analyst to join their team…
Remote
$82,000 - $95,000
A leading organization in the digital banking sector is seeking a talented Software Engineer to play a pivotal role in…