back to home

Best Open Source gpu Libraries

A curated list of the most popular GitHub repositories tagged with gpu. Select any project to visualize its architecture and dive into the codebase using RepoMind's AI engine.

#1pytorch/pytorch

Tensors and Dynamic neural networks in Python with strong GPU acceleration

98,347Python
Explore Repo

#2alacritty/alacritty

A cross-platform, OpenGL terminal emulator.

62,995Rust
Explore Repo

#3deepspeedai/DeepSpeed

DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.

41,835Python
Explore Repo

#4exelban/stats

macOS system monitor in your menu bar

37,091Swift
Explore Repo

#5taichi-dev/taichi

Productive, portable, and performant GPU programming in Python.

28,064C++
Explore Repo

#6fastai/fastai

The fastai deep learning library

27,914Jupyter Notebook
Explore Repo

#7Rem0o/FanControl.Releases

This is the release repository for Fan Control, a highly customizable fan controlling software for Windows.

18,987
Explore Repo

#8gfx-rs/wgpu

A cross-platform, safe, pure-Rust graphics API.

16,676Rust
Explore Repo

#9PavelDoGreat/WebGL-Fluid-Simulation

Play with fluids in your browser (works even on mobile)

16,183JavaScript
Explore Repo

#10gpujs/gpu.js

GPU Accelerated JavaScript

15,376JavaScript
Explore Repo

#11seerge/g-helper

Lightweight Armoury Crate alternative for Asus laptops with nearly the same functionality. Works with ROG Zephyrus, Flow, TUF, Strix, Scar, ProArt, Vivobook, Zenbook, Expertbook, ROG Ally, and many more.

14,964C#
Explore Repo

#12neovide/neovide

No Nonsense Neovim Client in Rust

14,823Rust
Explore Repo

#13pycaret/pycaret

An open-source, low-code machine learning library in Python

9,712Jupyter Notebook
Explore Repo

#14skypilot-org/skypilot

Run, manage, and scale AI workloads on any AI infrastructure. Use one system to access & manage all AI compute (Kubernetes, 20+ clouds, or on-prem).

9,582Python
Explore Repo

#15rapidsai/cudf

cuDF - GPU DataFrame Library

9,543C++
Explore Repo

#16NVIDIA/cutlass

CUDA Templates and Python DSLs for High-Performance Linear Algebra

9,447C++
Explore Repo

#17OlafenwaMoses/ImageAI

A python library built to empower developers to build applications and systems with self-contained Computer Vision capabilities

8,864Python
Explore Repo

#18catboost/catboost

A fast, scalable, high performance Gradient Boosting on Decision Trees library, used for ranking, classification, regression and other machine learning tasks for Python, R, Java, C++. Supports computation on CPU and GPU.

8,845C++
Explore Repo

#19ultralight-ux/Ultralight

Lightweight, high-performance HTML renderer for game and app developers.

4,945CMake
Explore Repo

#20ProjectPhysX/FluidX3D

The fastest and most memory efficient lattice Boltzmann CFD software, running on all GPUs and CPUs via OpenCL. Free for non-commercial use.

4,932C++
Explore Repo

#21arrayfire/arrayfire

ArrayFire: a general purpose GPU library.

4,868C++
Explore Repo

#22sktime/pytorch-forecasting

Time series forecasting with PyTorch

4,830Python
Explore Repo

#23MegEngine/MegEngine

MegEngine 是一个快速、可拓展、易于使用且支持自动求导的深度学习框架

4,807C++
Explore Repo

#24NVIDIA/cccl

CUDA Core Compute Libraries

2,495C++
Explore Repo

#25lupinemachines/lupine

LUPINE is a GPU over IP bridge allowing GPUs on remote machines to be attached to CPU-only machines.

2,419C++
Explore Repo

#26dstackai/dstack

Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.

2,229Python
Explore Repo

#27tenstorrent/tt-metal

:metal: TT-NN operator library, and TT-Metalium low level kernel programming model.

1,651C++
Explore Repo

#28NVIDIA/nvcf

Platform for deploying and routing GPU-accelerated inference, streaming, and batch workloads at scale.

203Go
Explore Repo

#29defilantech/LLMKube

Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.

200Go
Explore Repo

#30gogpu/wgpu

Pure Go WebGPU Implementation

174Go
Explore Repo