Data Engineer, Retrieval & Training Pipelines
Build the retrieval, preparation, and training pipelines that feed every model in Atlas.
Atlas runs on operational data: events, inventories, movements, sensor feeds, and the records that operators and institutions generate every day. This role owns the systems that find that data, retrieve it correctly, prepare it for training and inference, and keep it flowing reliably as the platform grows.
- Build and operate the retrieval layer: how Atlas finds and pulls operational data from diverse sources, with correct provenance and audit trail.
- Design and maintain the training data pipelines: ingestion, cleaning, feature engineering, and versioning for machine learning workloads.
- Own the schema and lineage of the data used for training. Models are only as trustworthy as the data behind them.
- Work with intelligence engineers on feature stores, offline and online consistency, and training-serving skew.
- Build tooling that makes data exploration and analysis fast for analysts and scientists.
- Strong data engineering background in production environments.
- Experience with analytical pipelines and ETL in production (dbt, Airflow, Spark, or equivalent).
- Strong SQL and comfort with at least one systems language (Python, Go, or similar).
- Attention to provenance and data quality. These are not afterthoughts for us.
These are not required. They are the kind of background that tends to do well in this role.
- Experience building retrieval or feature pipelines for machine learning systems.
- Experience with event streaming or event-sourced architectures.
- Geospatial data experience.
Apply for this role
Applications are read by the team that owns the role. We respond to every applicant. If you do not hear back within two weeks, write again.
Related roles
Applied Data Scientist, Operational Analytics
Find the best ways to analyze operational data and turn it into the analytical foundations of Atlas intelligence.
Senior Machine Learning Engineer, Decision Intelligence
Design, train, and operate the predictive models and fine-tuned systems that produce decision intelligence across operational domains.
Research Scientist, Simulation & Decision Systems
Lead the methodological foundations of simulation, scenario reasoning, and decision support in Atlas.
Bring connected intelligence online in weeks, not years.
We begin with a focused engagement on one operational domain, proving traceable, explainable intelligence end to end, then scale the architecture across the wider network.
Scoping
Define the domain, data sources, and decision owners.
Deployment
Connect signals, calibrate models, stand up provenance.
Operations
Live intelligence, weekly briefings, capability transfer.