← Explore
TOPIC

#inference-engine

Open source repositories tagged with #inference-engine, ranked by health score.

EfficientMoE
EfficientMoE/MoE-Infinity
Python
89
health

PyTorch library for cost-effective, fast and easy serving of MoE models.

★ 367
zhongkaifu
zhongkaifu/TensorSharp
C#
89
health

A native .NET LLM inference engine and agent runtime for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, iPhone App, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/iOS/Linux with full GPU capability

★ 557
ovg-project
ovg-project/kvcached
Python
88
health

Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond

★ 1.5k
carloslfu
carloslfu/slotstream
Swift
88
health

Run a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts in memory, so it runs on Macs with 16 to 64 GB. One native Swift binary on MLX and Metal, no Python, offline. Works with Claude Code, Codex and Ollama or OpenAI clients.

★ 423
headpiece747
headpiece747/ninfer-5090-windows
C++
84
health

Native Windows port of NInfer engine for RTX 5090. Features Qwen3.8-27B with QUASAR and NInfer models, MTP/DFlash2 with vision and 262,144 context

★ 55
ariannamethod
ariannamethod/notorch
C
82
health

neural networks in pure C

★ 27
prabanta-dev
prabanta-dev/janas
C
80
health

Janas-LLM: large mixture-of-experts language models on ordinary computers, in C. Experts streamed from NVMe, with or without a GPU.

★ 32