HOME/CAREERS/Data Pipeline Engineer

Data Pipeline Engineer

DATARemote · Full-timeApply for this role →

Models are only as good as the data feeding them — and the most expensive GPUs in the world shouldn't sit idle waiting on a pipeline. As a Data Pipeline Engineer you'll build the automated ingestion, transformation, and governance layer that keeps our systems, and our customers' models, fed with clean, production-ready data.

You'll own the data plane end to end: reliable, observable, and governed by default.

What you'll do

Build and operate automated, observable pipelines for ingesting data from any source.
Implement transformation, validation, and deduplication at scale, with data quality enforced in-pipeline.
Bake governance and lineage into every pipeline — so data is traceable, compliant, and trusted.
Optimise data delivery to keep training hardware saturated and never waiting on the pipeline.
Partner with strategy and deployment teams to stand up pipelines as part of customer deployments.
Own reliability: monitoring, alerting, and recovery for the data plane.

What we're looking for

4+ years in data engineering, building production pipelines at scale.
Strong with a modern data stack — orchestration (Airflow / Dagster / Prefect), distributed processing (Spark / Ray / Flink), and cloud or on-prem object storage.
Solid software-engineering fundamentals (Python and/or Scala/Go), testing, and CI/CD.
Experience with data quality, governance, and lineage tooling.
Comfortable owning systems end to end, from design through on-call.

Nice to have

Experience feeding large-scale ML / LLM training pipelines.
Familiarity with GPU data-loading bottlenecks and high-throughput storage.
Exposure to air-gapped or on-prem environments.
DEPARTMENT
DATA
ARRANGEMENT
Remote · Full-time
Apply for this role →← All open roles
Not quite the right fit?

Tell us what you'd want to build.

We're always looking for exceptional engineers. Send a note and a CV — we read every one.