catalog entry
LLMKube
Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Metal support, and OpenAI-compatible API.
signals
- stars
- 208
- commits 3mo
- 467
- last release
- v0.9.252026-09-05
- status
- active
Stars count attention, not fitness. The release date and the three-month commit count say more about whether anyone is still maintaining this.
No Wellworn verdict has been tested against LLMKube yet, so nothing on this page is a recommendation.
also tagged Generative Artificial Intelligence (GenAI)
- Ollama180,418 stars
- Open-WebUI151,260 stars
- LobeHub82,297 stars
- AnythingLLM65,748 stars
- LocalAI48,963 stars
- licence
- Apache-2.0
- platforms
- Go, Docker, K8S
- categories
- Generative Artificial Intelligence (GenAI)
Imported from awesome-selfhosted under CC BY-SA 3.0. Wellworn has not verified it.