back to home

Best Open Source cuda Libraries

A curated list of the most popular GitHub repositories tagged with cuda. Select any project to visualize its architecture and dive into the codebase using RepoMind's AI engine.

#1vllm-project/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

73,416Python
Explore Repo

#2sgl-project/sglang

SGLang is a high-performance serving framework for large language models and multimodal models.

32,971Python
Explore Repo

#3hashcat/hashcat

World's fastest and most advanced password recovery utility

25,601C
Explore Repo

#4NVlabs/instant-ngp

Instant neural graphics primitives: lightning fast NeRF and more

17,315Cuda
Explore Repo

#5kaldi-asr/kaldi

kaldi-asr/kaldi is the official location of the Kaldi project.

15,343Shell
Explore Repo

#6NVIDIA/TensorRT-LLM

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

14,511Python
Explore Repo

#7LMCache/LMCache

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

11,582Python
Explore Repo

#8xlite-dev/LeetCUDA

📚LeetCUDA: Modern CUDA Learn Notes with PyTorch for Beginners🐑, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.🎉

9,920Cuda
Explore Repo

#9rapidsai/cudf

cuDF - GPU DataFrame Library

9,543C++
Explore Repo

#10NVIDIA/cutlass

CUDA Templates and Python DSLs for High-Performance Linear Algebra

9,447C++
Explore Repo

#11Oneflow-Inc/oneflow

OneFlow is a deep learning framework designed to be user-friendly, scalable and efficient.

9,389C++
Explore Repo

#12replicate/cog

Containers for machine learning

9,268Go
Explore Repo

#13NVIDIA/cuda-samples

Samples for CUDA Developers which demonstrates features in CUDA Toolkit

8,961C
Explore Repo

#14catboost/catboost

A fast, scalable, high performance Gradient Boosting on Decision Trees library, used for ranking, classification, regression and other machine learning tasks for Python, R, Java, C++. Supports computation on CPU and GPU.

8,845C++
Explore Repo

#15gpustack/gpustack

A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.

5,573Python
Explore Repo

#16arrayfire/arrayfire

ArrayFire: a general purpose GPU library.

4,868C++
Explore Repo

#17NVIDIA/cccl

CUDA Core Compute Libraries

2,495C++
Explore Repo

#18lupinemachines/lupine

LUPINE is a GPU over IP bridge allowing GPUs on remote machines to be attached to CPU-only machines.

2,419C++
Explore Repo

#19tenstorrent/tt-metal

:metal: TT-NN operator library, and TT-Metalium low level kernel programming model.

1,651C++
Explore Repo

#20SemiAnalysisAI/InferenceX

Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3

1,591Python
Explore Repo

#21hybridgroup/yzma

Go with your own intelligence - Go applications that directly integrate llama.cpp for local inference using hardware acceleration.

574Go
Explore Repo

#22a2flo/floor

A C++ Compute/Graphics Library and Toolchain enabling same-source CUDA/Host/Metal/OpenCL/Vulkan C++ programming and execution.

341C++
Explore Repo

#23avifenesh/memra

Rust + CUDA inference engine for NVIDIA RTX PRO 6000 Blackwell and RTX 5090. Serves safetensors and GGUF over an OpenAI-compatible API, with per-device tuned defaults and speculative decode gated byte-identical to plain decode. Hosted instance: inference.tiyuvta.ai

328Rust
Explore Repo