← Explore
TOPIC

#speculative-decoding

Open source repositories tagged with #speculative-decoding, ranked by health score.

avifenesh
avifenesh/memra
Rust
89
health

Rust + CUDA inference engine for NVIDIA RTX PRO 6000 Blackwell and RTX 5090. Serves safetensors and GGUF over an OpenAI-compatible API, with per-device tuned defaults and speculative decode gated byte-identical to plain decode. Hosted instance: inference.tiyuvta.ai

328
Luce-Org
Luce-Org/lucebox
C++
88
health

LLM speculative inference server for heterogeneous hardware & consumer GPUs

2.8k
Avarok-Cybersecurity
Avarok-Cybersecurity/atlas
Rust
83
health

Pure Rust Inference Engine

672