← Explore
TOPIC

#cuda

Open source repositories tagged with #cuda, ranked by health score.

debpalash
debpalash/VoiceStudio
Python
90
health

VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

11.3k
avifenesh
avifenesh/memra
Rust
89
health

Rust + CUDA inference engine for NVIDIA RTX PRO 6000 Blackwell and RTX 5090. Serves safetensors and GGUF over an OpenAI-compatible API, with per-device tuned defaults and speculative decode gated byte-identical to plain decode. Hosted instance: inference.tiyuvta.ai

321
zhongkaifu
zhongkaifu/TensorSharp
C#
89
health

A native .NET LLM inference engine for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/Linux with full GPU capability

376
meta-pytorch
meta-pytorch/torchrec
Python
88
health

Pytorch domain library for recommendation systems

2.6k
NVIDIA
NVIDIA/warp
Python
88
health

A Python framework for GPU-accelerated simulation, robotics, and machine learning.

7.0k
NVIDIA
NVIDIA/TransformerEngine
Python
88
health

A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference.

3.5k
NVlabs
NVlabs/cuda-oxide
Rust
88
health

cuda-oxide is an experimental Rust-to-CUDA compiler that lets you write (SIMT) GPU kernels in safe(ish), idiomatic Rust. It compiles standard Rust code directly to PTX — no DSLs, no foreign language bindings, just Rust.

3.1k
Zaneham
Zaneham/Booth
C
88
health

Open-source CUDA, Triton and HIP compiler targeting multiple GPU and CPU architectures.

1.7k
luigifcruz
luigifcruz/CyberEther
C++
88
health

High-performance GPU-accelerated signal processing and visualization framework that runs anywhere.

780
lupinemachines
lupinemachines/lupine
C++
88
health

LUPINE is a GPU over IP bridge allowing GPUs on remote machines to be attached to CPU-only machines.

2.4k
invergent-ai
invergent-ai/surogate
C++
88
health

Training/Fine-tuning at the speed of light

812
uccl-project
uccl-project/uccl
C++
88
health

UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)

1.5k
MrNeRF
MrNeRF/LichtFeld-Studio
C++
88
health

Train, inspect, edit, automate, and export 3D Gaussian Splatting scenes from a single native application.

3.6k
pytorch
pytorch/TensorRT
Python
87
health

PyTorch/TorchScript/FX compiler for NVIDIA GPUs using TensorRT

3.0k
flashinfer-ai
flashinfer-ai/flashinfer
Python
87
health

FlashInfer: Kernel Library for LLM Serving

6.2k
Indras-Mirror
Indras-Mirror/llama.cpp-turboq-mtp
C++
87
health

Fused TBQ4 Flash Attention + MTP + Shared Tensors + Qwen35 SWA Hybrid for llama.cpp — 82+ tok/s, lossless 4.25 bpv KV cache, SWA-bounded deep-context decode (w/ long-range recall) on RTX 4090

90
sonots
sonots/cumo
C
86
health

Cumo (pronounced like "koomo") is CUDA aware numerical library whose interface is highly compatible with Ruby Numo

99
NVIDIA
NVIDIA/cudf
C++
86
health

cuDF - GPU DataFrame Library

9.7k
Avarok-Cybersecurity
Avarok-Cybersecurity/atlas
Rust
82
health

Pure Rust Inference Engine

658
iree-org
iree-org/iree
C++
80
health

A retargetable MLIR-based machine learning compiler and runtime toolkit.

3.9k