← Explore
TOPIC

#local-llm

Open source repositories tagged with #local-llm, ranked by health score.

nicedreamzapp
nicedreamzapp/claude-code-local
Python
90
health

Run Claude Code 100% on-device with local AI on Apple Silicon. MLX-native Anthropic-API server. 6 fighters incl. Muse-Glimmer 30B (now multimodal — reads images, abliterated), Gemma 4 31B, Qwen 3.5 122B (65 tok/s), DeepSeek V4 Flash (1M ctx). Private, offline, airgap-ready. Built for NDA / legal / healthcare workflows.

★ 3.3k
orneryd
orneryd/NornicDB
Go
88
health

Nornicdb is a distributed low-latency, Graph+Vector, Temporal MVCC with all sub-ms HNSW search, graph traversal, and writes. Using Neo4j Bolt/Cypher and qdrant's gRPC means you can switch with no changes while adding intelligent features like schemas, managed embeddings, reranking+llm, GPU accel, Auto-TLP, Policy-based Memory Decay, and MCP server.

★ 897
carloslfu
carloslfu/slotstream
Swift
88
health

Run a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts in memory, so it runs on Macs with 16 to 64 GB. One native Swift binary on MLX and Metal, no Python, offline. Works with Claude Code, Codex and Ollama or OpenAI clients.

★ 423
Nehanth
Nehanth/pooled
JavaScript
87
health

Peer-to-peer LLM inference in the browser: pool your devices to run big open models, for chat and coding agents.

★ 576
mindroom-ai
mindroom-ai/mindroom
Python
87
health

AI agents that know you and your work, in a chat app anyone can use. Open source, any model, self-host or hosted.

★ 320
LearningCircuit
LearningCircuit/local-deep-research
Python
86
health

~95% on SimpleQA (e.g. Qwen3.6-27B on a 3090). Supports all local and cloud LLMs (llama.cpp, Ollama, Google, ...). 10+ search engines - arXiv, PubMed, your private documents. Everything Local & Encrypted.

★ 9.2k
enapt
enapt/SwarmLLM
Rust
85
health

Free, open-source app that runs AI chat models on your own computer — and lets computers team up over the internet to run models too big for one machine. No account, no crypto. OpenAI- and Anthropic-compatible API.

★ 58
darshi1337
darshi1337/apogee
JavaScript
84
health

Private AI summarizer for anything you read. Inspired by Mozilla's Orbit.

★ 70
prabanta-dev
prabanta-dev/janas
C
80
health

Janas-LLM: large mixture-of-experts language models on ordinary computers, in C. Experts streamed from NVMe, with or without a GPU.

★ 32
leehack
leehack/llamadart
Dart
77
health

Cross-platform on-device LLM inference for Dart & Flutter: GGUF via llama.cpp and LiteRT-LM on Android, iOS, macOS, Windows, Linux and web.

★ 78
anthony-chaudhary
anthony-chaudhary/fak
Go
73
health

Agentic Runtime

★ 41
hertz-ai
hertz-ai/HARTOS
Python
72
health

An AI-native OS. Models run on your own hardware, nodes federate peer-to-peer with no broker, and the API is OpenAI-compatible. Boots, has its own Wayland compositor, and runs on 8GB. Apache 2.0.

★ 51
benmaster82
benmaster82/writher
Python
62
health

🎙️ Offline voice productivity for Windows - dictate text into any app and control a local AI assistant by voice. Manages notes, to-do lists, appointments & reminders. 100% local: faster-whisper + Ollama + SQLite. No cloud, no telemetry. Free and open source alternative to Wispr Flow

★ 86