back to home

Best Open Source data engineering Libraries

A curated list of the most popular GitHub repositories tagged with data engineering. Select any project to visualize its architecture and dive into the codebase using RepoMind's AI engine.

#1apache/superset

Apache Superset is a Data Visualization and Data Exploration Platform

70,995TypeScript
Explore Repo

#2GokuMohandas/Made-With-ML

Learn how to develop, deploy and iterate on production-grade ML applications.

46,810Jupyter Notebook
Explore Repo

#3apache/airflow

Apache Airflow - A platform to programmatically author, schedule, and monitor workflows

44,671Python
Explore Repo

#4DataTalksClub/data-engineering-zoomcamp

Data Engineering Zoomcamp is a free 9-week course on building production-ready data pipelines. The next cohort starts in January 2026. Join the course here 👇🏼

39,140Jupyter Notebook
Explore Repo

#5eugeneyan/applied-ml

📚 Papers & tech blogs by companies sharing their work on data science & machine learning in production.

28,715
Explore Repo

#6Avaiga/taipy

Turns Data and AI algorithms into production-ready web applications in no time.

19,113Python
Explore Repo

#7argoproj/argo-workflows

Workflow Engine for Kubernetes

16,531Go
Explore Repo

#8dagster-io/dagster

An orchestration platform for the development, production, and observation of data assets.

15,112Python
Explore Repo

#9andkret/Cookbook

The Data Engineering Cookbook

14,995Python
Explore Repo

#10semantica-agi/semantica

Graph-Native Infrastructure for Context and Accountable AI Systems

10,496Python
Explore Repo

#11xonsh/xonsh

🐚 Python-powered shell. Full-featured, cross-platform and AI-friendly.

9,251Python
Explore Repo

#12risingwavelabs/risingwave

Event streaming platform for agents, apps, and analytics. Continuously ingest, transform, and serve event data in real time, at scale.

8,862Rust
Explore Repo

#13mage-ai/mage-ai

🧙 Build, run, and manage data pipelines for integrating and transforming data.

8,677Python
Explore Repo

#14redpanda-data/connect

Fancy stream processing made operationally mundane

8,608Go
Explore Repo

#15lakehq/sail

Drop-in Apache Spark replacement written in Rust, unifying batch processing, stream processing, and compute-intensive AI workloads.

3,325Rust
Explore Repo

#16Zleap-AI/SAG

A new SOTA for RAG — an original retrieval architecture and an open-source knowledge base for humans and agents.

2,409Python
Explore Repo

#17DataRecce/recce

The data-validation toolkit for enhanced dbt (data build tool) PR review

472TypeScript
Explore Repo

#18snowflakedb/snowpark-python

Snowflake Snowpark Python API

340Python
Explore Repo

#19Hebbian-Robotics/hflow

Open source SDK for building multimodal data-quality, processing, enrichment, and curation pipelines for robotics and Physical AI.

123Python
Explore Repo

#20animetanver-commits/fabric-unified-data-blueprint

Microsoft Fabric Unified Data Foundation with Databricks & Purview 2026

115HTML
Explore Repo