Open source repositories tagged with #inference, ranked by health score.
Open-source library of optimized deep learning operations (matmul, convolution, attention) for CPUs (x64, AArch64, RISC-V) and Intel GPUs. Used by PyTorch, TensorFlow, OpenVINO, and ONNX Runtime.
OpenVINO™ is an open source toolkit for optimizing and deploying AI inference
Make videos on the computer you already own. FreeVideo runs MiniMax H3 in as little as 8 GB of VRAM and 16 GB of RAM, and adapts its acceleration path to your hardware.
Camelid: a Rust-native local inference backend with evidence-gated model compatibility.
⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache compression, MACOS + iOS iPhone app.
llm-d Router: The intelligent entry point for inference requests
A programmable Mixture-of-Models router for heterogeneous LLM inference
Community maintained hardware plugin for vLLM on Huawei Ascend
Distributed LLM inference fabric for heterogeneous local clusters. Features zero-copy SPSC ring buffers, automated tuning, and an OpenAI-compatible API.
A framework for efficient model inference with omni-modality models
Runs openpilot's large driving model on Jetson Orin Nano, macOS, iOS, Android, CUDA laptops with only your Comma, and a USB3 cable
From PyTorch model to end-to-end TensorRT inference experience in two commands—AI-native, cross-platform, and built for the best possible user experience.