Open source repositories tagged with #cuda, ranked by health score.
A native .NET LLM inference engine and agent runtime for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, iPhone App, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/iOS/Linux with full GPU capability
Local AI song generator with an editable score — YuE2 on your GPU: full songs with vocals, sheet music, covers, exact replay. Native Windows app, no Python, installer with auto-update.
NumPy & SciPy for GPU
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
NVIDIA's CUDA platform for Rust. Host runtime crates plus Tile (cutile-rs) and SIMT (cuda-oxide) kernel programming models in idiomatic Rust.
Train and serve LLMs at extreme speed and massive throughput.
Real-time multimodal desktop agent evolving toward a persistent AI OS interface (0.15 α).
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
PyTorch/TorchScript/FX compiler for NVIDIA GPUs using TensorRT
LuxCore source repository
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Open Source Inference Research Platform Standard / 开源推理研究平台
Native Windows port of NInfer engine for RTX 5090. Features Qwen3.8-27B with QUASAR and NInfer models, MTP/DFlash2 with vision and 262,144 context
neural networks in pure C
ONNX Runtime Server: The ONNX Runtime Server is a server that provides TCP and HTTP/HTTPS REST APIs for ONNX inference.
KeepGPU is a simple CLI app that keeps your GPUs running.