Open source repositories tagged with #llm-serving, ranked by health score.
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Open Source Inference Research Platform Standard / 开源推理研究平台
Agentic Runtime