Run GGUF language models locally with
llama.cpp,
CPU or GPU acceleration, and
llama-swap.
The server exposes an OpenAI-compatible API and discovers models placed in
the configured install directory, which defaults to ~/ai/models.
If you find Local AI Server useful, you can support its continued development:
⭐ Starring and sharing the repository also helps a lot.
| Feature | Local AI Server | Ollama | LM Studio | OpenAI / Gemini |
|---|---|---|---|---|
| Runs fully locally and privately | ✅ | ✅ | ✅ | ❌ |
| Designed for Linux servers | ✅ | ✅ | ⚠️ Desktop-focused | ❌ |
| Uses your existing GGUF files directly | ✅ | ⚠️ Import required | ✅ | ❌ |
| Automatic multi-model switching | ✅ | ✅ | ✅ | Cloud-managed |
| OpenAI-compatible API | ✅ | ✅ | ✅ | ✅ |
| User-level systemd service | ✅ | ⚠️ Usually system-wide | ❌ | Not applicable |
Transparent llama.cpp configuration |
✅ | ⚠️ Abstracted | ⚠️ GUI-managed | ❌ |
| No API fees | ✅ | ✅ | ✅ | ❌ |
| Pure Bash, no extra runtime | ✅ | ⚠️ Ships Go binary | ⚠️ Electron app | ❌ |
| Readable scripts, easy to audit | ✅ | ⚠️ Compiled binary | ⚠️ GUI app | ❌ |
| Light footprint, fits minimal VPS | ✅ | ⚠️ Moderate | ⚠️ Heavy desktop | ❌ |
Local AI Server gives Linux users a private, lightweight and transparent way to run multiple GGUF models through one OpenAI-compatible API, with automatic model switching and systemd service management — all in readable Bash scripts with no runtime to install.
.gguf model files, with per-model auto-tuning for GPU hardwarelocalai command for service, model, update, and uninstall taskssudo access during installationThe installer downloads current llama.cpp and llama-swap releases and can
install required packages with apt-get, dnf, or yum.
One-line install:
curl -fsSL https://hossbit.github.io/localai/install.sh | bash
CPU-only install:
curl -fsSL https://hossbit.github.io/localai/install.sh | LLAMA_CPP_BACKEND=cpu bash
Already installed and want another backend, including CUDA? No need to re-run the installer:
localai backend install vulkan
localai backend install rocm
localai backend install cuda
The default install directory is ~/ai. See the wiki’s Install
LocalAI
page for custom directories, manual installs, backend selection, and pinned
component versions.
By default LocalAI tracks upstream llama.cpp’s b[NUM] bleeding-edge builds
(cut on nearly every commit). To track its slower vX.Y.Z stable releases
instead, set LLAMA_CPP_CHANNEL=stable (see localai.conf).
Besides the prebuilt CPU/Vulkan/ROCm/OpenVINO/SYCL backends, LocalAI can also
build llama.cpp from source against your own NVIDIA CUDA Toolkit. Once
LocalAI is installed, add or switch backends with the CLI – no need to
re-run the installer, and no redownload or rebuild once a backend has been
installed before:
localai backend install cuda # build CUDA without switching to it yet
localai switch cuda # switch to it (installs first if needed)
localai switch vulkan # switch back -- instant
localai backend list # see what's installed and which is active
See the wiki’s Backend Selection
page for CUDA prerequisites (nvcc vs. nvidia-smi), choosing a backend
during the initial install, tuning overrides, and Fedora/RHEL notes.
LocalAI discovers GGUF files from:
~/ai/models
Download a .gguf model from a source such as Hugging Face, then put it in that
directory. After adding or removing models, reload LocalAI so it regenerates
the config and restarts only if it changed, then list the detected models:
localai reload
localai models
For a single-file model, either place the file directly in ~/ai/models:
~/ai/models/Qwen2.5-Coder-7B-Instruct-Q4_K_M.gguf
or keep it in its own folder:
~/ai/models/Qwen2.5-Coder-7B-Instruct-Q4_K_M/
`-- Qwen2.5-Coder-7B-Instruct-Q4_K_M.gguf
For split GGUF models, keep all shards together in one folder. The first shard
must use canonical llama.cpp split naming, such as 00001-of-00003:
~/ai/models/DeepSeek-V4-Flash-UD-IQ1_M/
|-- DeepSeek-V4-Flash-UD-IQ1_M-00001-of-00003.gguf
|-- DeepSeek-V4-Flash-UD-IQ1_M-00002-of-00003.gguf
`-- DeepSeek-V4-Flash-UD-IQ1_M-00003-of-00003.gguf
LocalAI registers only the first shard. llama.cpp loads the remaining shards
automatically.
Recommended layout:
~/ai/models/
|-- Qwen2.5-Coder-7B-Instruct-Q4_K_M.gguf
|-- Mistral-7B-Instruct-Q4_K_M/
| `-- Mistral-7B-Instruct-Q4_K_M.gguf
`-- DeepSeek-V4-Flash-UD-IQ1_M/
|-- DeepSeek-V4-Flash-UD-IQ1_M-00001-of-00003.gguf
|-- DeepSeek-V4-Flash-UD-IQ1_M-00002-of-00003.gguf
`-- DeepSeek-V4-Flash-UD-IQ1_M-00003-of-00003.gguf
If LocalAI warns that files look like non-canonical split fragments, rename the files to llama.cpp split format or merge them first:
llama-gguf-split --merge first-fragment.gguf merged-model.gguf
Use localai suggest after adding large models to get advisory runtime settings
based on your installed models, RAM, backend, and detected GPU memory. It uses
the actual GGUF file size as the base estimate, not an exact parameter-count
formula. Runtime memory also depends on context length, KV cache type, batch
size, backend buffers, and operating-system headroom.
GPU-backed installs auto-tune per-model GPU layers, KV cache type, and
flash-attention from your hardware, and enable free self-speculative decoding
by default. On multi-GPU systems, LOCALAI_SPLIT_MODE, LOCALAI_TENSOR_SPLIT,
LOCALAI_MAIN_GPU, and LOCALAI_DEVICE control how models are placed across
devices. See the wiki for per-model overrides (models.d), multi-GPU tuning,
MoE CPU offload, reasoning-model tuning, multimodal --mmproj setup,
speculative-decoding tuning, LoRA adapters, metrics, and startup preloading.
Start LocalAI:
localai start
localai check
The API is available at http://127.0.0.1:$(cat ~/ai/conf/port)/v1.
Two browser UIs are available with nothing extra to install — llama-swap and llama.cpp both ship one, and LocalAI just wires them up:
http://127.0.0.1:PORT/ui. Shows model status,
lets you load/unload models, and tails logs.http://127.0.0.1:PORT/upstream/MODEL_ID/. The
same chat UI llama-server serves standalone, proxied per model through
llama-swap (loading the model on first request, like any other API call).Run localai ui to print both URLs, or localai ui MODEL_ID for one
model’s chat UI directly:
$ localai ui
llama-swap Web UI (model status, load/unload, logs):
http://127.0.0.1:11435/ui
llama.cpp chat UI for a specific model:
http://127.0.0.1:11435/upstream/MODEL_ID/
(run 'localai ui MODEL_ID', or 'localai models' for exact IDs)
API-key auth is enabled. When the browser prompts for credentials,
leave the username blank and use your saved API key as the password.
Need a new key? Run 'localai key create browser'; the secret is shown once.
$ localai ui Qwen2.5-Coder-7B-Instruct-Q4_K_M
llama.cpp chat UI for Qwen2.5-Coder-7B-Instruct-Q4_K_M:
http://127.0.0.1:11435/upstream/Qwen2.5-Coder-7B-Instruct-Q4_K_M/
Use localai ui --open to launch the dashboard in your desktop browser, or
localai ui --open MODEL_ID for a model’s chat UI. On SSH/headless systems,
use the printed URL from a browser that can reach the server. The dashboard
in llama-swap v257 includes a searchable model picker, responsive chat, and
live generation statistics.
localai key list shows key metadata, not the secret. Use your saved key or
create a new one with localai key create browser.
Open the printed URL in a browser. If API-key auth is enabled (API keys), the browser’s login prompt wants the username left blank and an active key as the password.
Most users only need these:
| Command | Purpose |
|---|---|
localai start |
Start the service. |
localai stop |
Unload loaded models, then stop the service. |
localai restart |
Restart the service. |
localai reload |
Rescan models and restart only if config.yaml would change; prints added/removed models. |
localai status |
Show service, process, API, and port status. |
localai check |
Check the API and model list. |
localai models |
List installed .gguf models and show loaded state when the API is reachable. |
localai ui [--open] [MODEL] |
Print a dashboard or model chat URL; optionally open it in your desktop browser. |
localai suggest |
Suggest runtime settings from installed model sizes and detected hardware. |
localai load MODEL |
Warm one model. |
localai unload MODEL |
Release one loaded model. |
localai key ... |
Manage API keys (create, list, revoke, rotate) — see API keys. |
localai update |
Update installed components. |
localai version |
Show component versions. |
localai uninstall |
Remove helper files; models are kept by default. |
By default the API has no authentication, matching llama-swap’s own default —
fine as long as LOCALAI_LISTEN_HOST stays 127.0.0.1 (loopback only, the
default). If you plan to reach it from another machine on your LAN, create at
least one key first.
localai key create work-laptop # name is just a label; shown once, then masked
localai key list # id, name, created, status, masked fingerprint
localai key revoke <id> # deactivate a key immediately
localai key rotate <id> # issue a replacement, then revoke the old one
create and rotate print the full secret exactly once, right after it’s
active — save it immediately, it cannot be shown again:
Key created: work-laptop
id: a1b2c3d4e5f6
created: 2026-07-23T16:32:10Z
Save this key now - it will not be shown again:
sk-localai-9f1a2b3c4d5e6f7089...
Use it as a Bearer token:
curl http://127.0.0.1:11435/v1/models \
-H "Authorization: Bearer sk-localai-9f1a2b3c4d5e6f7089..."
Once at least one key is active, every request needs it:
curl http://127.0.0.1:11435/v1/models \
-H "Authorization: Bearer sk-localai-REPLACE_ME"
from openai import OpenAI
client = OpenAI(
base_url="http://127.0.0.1:11435/v1",
api_key="sk-localai-REPLACE_ME",
)
Behavior notes:
localai key create).Authorization: Bearer
header.LOCALAI_REQUIRE_API_KEY=1 in localai.conf, which makes config
generation refuse to produce an unauthenticated config at all.rotate the key (or revoke it and
create a new one).conf/api-keys.tsv (mode 0600, outside Git). Active keys
are also rendered into conf/keys.d/keys.yaml (mode 0600), each with a
# name comment above it, and merged into the running config via
llama-swap’s --config-dir — config.yaml itself never contains a
plaintext key. Manage keys through localai key ..., not by hand-editing
either file; keys.yaml is removed entirely once no keys are active.localai reload already uses for models).API keys authenticate requests; they don’t encrypt them. For LAN/WAN access, put a TLS-terminating reverse proxy in front and restrict it with a firewall — see Security.
The helper scripts bind llama-swap to 127.0.0.1, so the API is available only
on the local machine by default. Do not expose it to a network without adding
authentication (localai key create), TLS, and appropriate
firewall rules.
This project is built on top of:
Special thanks to the maintainers and contributors of these projects.
LocalAI focuses on simplifying installation, configuration, model management, and service deployment for local LLM environments.