llama.cpp Server, an efficient local inference engine for open LLMs, ready to serve models from first boot.
llama.cpp is a high efficiency inference engine that runs open large language models on CPU and GPU hardware, exposing an OpenAI compatible server. It is used to host and serve local LLMs with minimal overhead.
cloudimg ships the llama.cpp server hardened and fully patched on Ubuntu 24.04 LTS, preconfigured so the inference endpoint is ready on first boot. 24/7 support is included with every image.
Real screenshots taken while testing this image against its deployment guide.