wellworn
catalog entry

LLMKube

Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Metal support, and OpenAI-compatible API.

website llmkube.comsource github.com
signals
stars
208
commits 3mo
467
last release
v0.9.252026-09-05
status
active

Stars count attention, not fitness. The release date and the three-month commit count say more about whether anyone is still maintaining this.

No Wellworn verdict has been tested against LLMKube yet, so nothing on this page is a recommendation.

also tagged Generative Artificial Intelligence (GenAI)

All Generative Artificial Intelligence (GenAI) tools

licence
Apache-2.0
platforms
Go, Docker, K8S
categories
Generative Artificial Intelligence (GenAI)

Imported from awesome-selfhosted under CC BY-SA 3.0. Wellworn has not verified it.

All categories