← Explore
TOPIC

#llm-inference

Open source repositories tagged with #llm-inference, ranked by health score.

openvinotoolkit
openvinotoolkit/openvino
C++
90
health

OpenVINO™ is an open source toolkit for optimizing and deploying AI inference

★ 11.0k
StayLameBro
StayLameBro/backburner
Python
89
health

Your iPhone helps your Mac run a 27B model: faster prompt reading and more context over a USB-C cable

★ 674
neuron-core
neuron-core/neuron-ai
PHP
89
health

The Agentic Framework of the PHP ecosystem to build production-ready AI driven applications. Connect components (LLMs, Tools, vector DBs, memory) to agents that interact with your data and UI.

★ 2.1k
incoai
incoai/splash
C++
88
health

A local inference engine for Apple silicon, built around the model.

★ 1.2k
carloslfu
carloslfu/slotstream
Swift
88
health

Run a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts in memory, so it runs on Macs with 16 to 64 GB. One native Swift binary on MLX and Metal, no Python, offline. Works with Claude Code, Codex and Ollama or OpenAI clients.

★ 420
lemonade-sdk
lemonade-sdk/lemonade
C++
86
health

Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk

★ 5.8k
kubernetes-sigs
kubernetes-sigs/lws
Go
84
health

LeaderWorkerSet: An API for deploying a group of pods as a unit of replication

★ 848
spiceai
spiceai/spiceai
Rust
81
health

Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query, search, and LLM-inference engine in Rust for data-grounded AI apps and agents.

★ 3.1k
townsendmerino
townsendmerino/goinfer
Go
81
health

Pure-Go, no-cgo local LLM inference — run Gemma, Qwen, Llama and friends from safetensors or GGUF in a single static binary.

★ 18