Atlas · skill

Kedro

Kedro is a Python framework for structuring data-science and machine-learning projects as explicit pipelines. It separates processing functions, data interfaces and configuration, helping teams turn exploratory work into repeatable components whose dependencies and outputs can be understood and tested.

toolETL/ELT

What it is

Kedro organizes computation into nodes with declared inputs and outputs, then connects nodes into pipelines. A data catalog defines how named datasets are loaded or saved, reducing the need to embed storage details in processing logic. Configuration separates environment-specific choices from the code that performs transformations. This structure makes dependency relationships visible and supports reuse, but it does not make arbitrary node code deterministic or correct. Kedro is a project and pipeline framework rather than a universal production scheduler; deployment and operational supervision depend on the surrounding tools and infrastructure chosen by the team.

What the work involves

The practitioner extracts reusable transformations from notebooks into functions, defines their data dependencies and configures catalog entries for storage. They test nodes independently and verify full pipeline behavior with controlled inputs. Useful artifacts include a pipeline graph and configuration that identifies each dataset interface. Sensitive credentials are handled through appropriate mechanisms rather than hard-coded into nodes. The team also decides how pipeline outputs are versioned and deployed, checking that environment changes preserve semantics instead of assuming separation of configuration eliminates every source of irreproducibility.

Illustrative example

A forecasting project has a notebook that reads files, cleans data and trains a model in one execution state. The engineer converts these stages into Kedro nodes and names their intermediate datasets in the catalog. Tests now exercise cleaning independently, while a full run produces a model from explicit inputs. A production environment changes storage locations through configuration without rewriting the transformation functions, and run records preserve which inputs were used.

Limits and common mistakes

A tidy pipeline graph can still contain hidden global state, nondeterminism or poorly defined data semantics. Catalog entries are not complete lineage or governance by themselves, and production operations require further choices. Kedro introduces conventions that should justify their maintenance cost. The useful test is whether another team member can run and change the pipeline with clear inputs and outputs, rather than whether every notebook has been mechanically converted into a node.

Prerequisites

No prerequisites.

Related skills

Sources and further reading

  • Kedro documentation

    Official pipeline, node and data-catalog concepts for structured data-science projects.

Last updated: 2026-10-10