local-ai-server

Local AI Server for Linux

LocalAI local LLM server
![Linux](https://img.shields.io/badge/Linux-x86--64-111827) ![Debian-based](https://img.shields.io/badge/Debian--based-apt--get-A81D33) ![Red Hat-based](https://img.shields.io/badge/Red%20Hat--based-dnf%20%7C%20yum-EE0000) ![llama.cpp](https://img.shields.io/badge/Engine-llama.cpp-6C5CE7) ![CUDA](https://img.shields.io/badge/CUDA-supported-76B900) ![API](https://img.shields.io/badge/API-OpenAI--compatible-111827) ![Service](https://img.shields.io/badge/Service-systemd%20user-F59E0B) ![License](https://img.shields.io/badge/License-MIT-10B981)

Run GGUF language models locally with llama.cpp, CPU or GPU acceleration, and llama-swap. The server exposes an OpenAI-compatible API and discovers models placed in the configured install directory, which defaults to ~/ai/models.

Latest Release Downloads GitHub Stars

❤️ Support

If you find Local AI Server useful, you can support its continued development:

Buy me a coffee

⭐ Starring and sharing the repository also helps a lot.

Why Local AI Server?

Feature Local AI Server Ollama LM Studio OpenAI / Gemini
Runs fully locally and privately ✅ ✅ ✅ ❌
Designed for Linux servers ✅ ✅ ⚠️ Desktop-focused ❌
Uses your existing GGUF files directly ✅ ⚠️ Import required ✅ ❌
Automatic multi-model switching ✅ ✅ ✅ Cloud-managed
OpenAI-compatible API ✅ ✅ ✅ ✅
User-level systemd service ✅ ⚠️ Usually system-wide ❌ Not applicable
Transparent llama.cpp configuration ✅ ⚠️ Abstracted ⚠️ GUI-managed ❌
No API fees ✅ ✅ ✅ ❌
Pure Bash, no extra runtime ✅ ⚠️ Ships Go binary ⚠️ Electron app ❌
Readable scripts, easy to audit ✅ ⚠️ Compiled binary ⚠️ GUI app ❌
Light footprint, fits minimal VPS ✅ ⚠️ Moderate ⚠️ Heavy desktop ❌

Main Advantage

Local AI Server gives Linux users a private, lightweight and transparent way to run multiple GGUF models through one OpenAI-compatible API, with automatic model switching and systemd service management — all in readable Bash scripts with no runtime to install.

What it provides

Requirements

The installer downloads current llama.cpp and llama-swap releases and can install required packages with apt-get, dnf, or yum.

Install

One-line install:

curl -fsSL https://hossbit.github.io/localai/install.sh | bash

CPU-only install:

curl -fsSL https://hossbit.github.io/localai/install.sh | LLAMA_CPP_BACKEND=cpu bash

Already installed and want another backend, including CUDA? No need to re-run the installer:

localai backend install vulkan
localai backend install rocm
localai backend install cuda

The default install directory is ~/ai. See the wiki’s Install LocalAI page for custom directories, manual installs, backend selection, and pinned component versions.

By default LocalAI tracks upstream llama.cpp’s b[NUM] bleeding-edge builds (cut on nearly every commit). To track its slower vX.Y.Z stable releases instead, set LLAMA_CPP_CHANNEL=stable (see localai.conf).

CUDA and Switching Backends

Besides the prebuilt CPU/Vulkan/ROCm/OpenVINO/SYCL backends, LocalAI can also build llama.cpp from source against your own NVIDIA CUDA Toolkit. Once LocalAI is installed, add or switch backends with the CLI – no need to re-run the installer, and no redownload or rebuild once a backend has been installed before:

localai backend install cuda   # build CUDA without switching to it yet
localai switch cuda            # switch to it (installs first if needed)
localai switch vulkan          # switch back -- instant
localai backend list           # see what's installed and which is active

See the wiki’s Backend Selection page for CUDA prerequisites (nvcc vs. nvidia-smi), choosing a backend during the initial install, tuning overrides, and Fedora/RHEL notes.

Add a model

LocalAI discovers GGUF files from:

~/ai/models

Download a .gguf model from a source such as Hugging Face, then put it in that directory. After adding or removing models, reload LocalAI so it regenerates the config and restarts only if it changed, then list the detected models:

localai reload
localai models

For a single-file model, either place the file directly in ~/ai/models:

~/ai/models/Qwen2.5-Coder-7B-Instruct-Q4_K_M.gguf

or keep it in its own folder:

~/ai/models/Qwen2.5-Coder-7B-Instruct-Q4_K_M/
`-- Qwen2.5-Coder-7B-Instruct-Q4_K_M.gguf

For split GGUF models, keep all shards together in one folder. The first shard must use canonical llama.cpp split naming, such as 00001-of-00003:

~/ai/models/DeepSeek-V4-Flash-UD-IQ1_M/
|-- DeepSeek-V4-Flash-UD-IQ1_M-00001-of-00003.gguf
|-- DeepSeek-V4-Flash-UD-IQ1_M-00002-of-00003.gguf
`-- DeepSeek-V4-Flash-UD-IQ1_M-00003-of-00003.gguf

LocalAI registers only the first shard. llama.cpp loads the remaining shards automatically.

Recommended layout:

~/ai/models/
|-- Qwen2.5-Coder-7B-Instruct-Q4_K_M.gguf
|-- Mistral-7B-Instruct-Q4_K_M/
|   `-- Mistral-7B-Instruct-Q4_K_M.gguf
`-- DeepSeek-V4-Flash-UD-IQ1_M/
    |-- DeepSeek-V4-Flash-UD-IQ1_M-00001-of-00003.gguf
    |-- DeepSeek-V4-Flash-UD-IQ1_M-00002-of-00003.gguf
    `-- DeepSeek-V4-Flash-UD-IQ1_M-00003-of-00003.gguf

If LocalAI warns that files look like non-canonical split fragments, rename the files to llama.cpp split format or merge them first:

llama-gguf-split --merge first-fragment.gguf merged-model.gguf

Use localai suggest after adding large models to get advisory runtime settings based on your installed models, RAM, backend, and detected GPU memory. It uses the actual GGUF file size as the base estimate, not an exact parameter-count formula. Runtime memory also depends on context length, KV cache type, batch size, backend buffers, and operating-system headroom.

GPU-backed installs auto-tune per-model GPU layers, KV cache type, and flash-attention from your hardware, and enable free self-speculative decoding by default. On multi-GPU systems, LOCALAI_SPLIT_MODE, LOCALAI_TENSOR_SPLIT, LOCALAI_MAIN_GPU, and LOCALAI_DEVICE control how models are placed across devices. See the wiki for per-model overrides (models.d), multi-GPU tuning, MoE CPU offload, reasoning-model tuning, multimodal --mmproj setup, speculative-decoding tuning, LoRA adapters, metrics, and startup preloading.

Use the server

Start LocalAI:

localai start
localai check

The API is available at http://127.0.0.1:$(cat ~/ai/conf/port)/v1.

Web UIs

Two browser UIs are available with nothing extra to install — llama-swap and llama.cpp both ship one, and LocalAI just wires them up:

Run localai ui to print both URLs, or localai ui MODEL_ID for one model’s chat UI directly:

$ localai ui
llama-swap Web UI (model status, load/unload, logs):
  http://127.0.0.1:11435/ui

llama.cpp chat UI for a specific model:
  http://127.0.0.1:11435/upstream/MODEL_ID/
  (run 'localai ui MODEL_ID', or 'localai models' for exact IDs)

API-key auth is enabled. When the browser prompts for credentials,
leave the username blank and use your saved API key as the password.
Need a new key? Run 'localai key create browser'; the secret is shown once.

$ localai ui Qwen2.5-Coder-7B-Instruct-Q4_K_M
llama.cpp chat UI for Qwen2.5-Coder-7B-Instruct-Q4_K_M:
  http://127.0.0.1:11435/upstream/Qwen2.5-Coder-7B-Instruct-Q4_K_M/

Use localai ui --open to launch the dashboard in your desktop browser, or localai ui --open MODEL_ID for a model’s chat UI. On SSH/headless systems, use the printed URL from a browser that can reach the server. The dashboard in llama-swap v257 includes a searchable model picker, responsive chat, and live generation statistics.

localai key list shows key metadata, not the secret. Use your saved key or create a new one with localai key create browser.

Open the printed URL in a browser. If API-key auth is enabled (API keys), the browser’s login prompt wants the username left blank and an active key as the password.

Service and helper commands

Most users only need these:

Command Purpose
localai start Start the service.
localai stop Unload loaded models, then stop the service.
localai restart Restart the service.
localai reload Rescan models and restart only if config.yaml would change; prints added/removed models.
localai status Show service, process, API, and port status.
localai check Check the API and model list.
localai models List installed .gguf models and show loaded state when the API is reachable.
localai ui [--open] [MODEL] Print a dashboard or model chat URL; optionally open it in your desktop browser.
localai suggest Suggest runtime settings from installed model sizes and detected hardware.
localai load MODEL Warm one model.
localai unload MODEL Release one loaded model.
localai key ... Manage API keys (create, list, revoke, rotate) — see API keys.
localai update Update installed components.
localai version Show component versions.
localai uninstall Remove helper files; models are kept by default.

API keys

By default the API has no authentication, matching llama-swap’s own default — fine as long as LOCALAI_LISTEN_HOST stays 127.0.0.1 (loopback only, the default). If you plan to reach it from another machine on your LAN, create at least one key first.

localai key create work-laptop   # name is just a label; shown once, then masked
localai key list                 # id, name, created, status, masked fingerprint
localai key revoke <id>          # deactivate a key immediately
localai key rotate <id>          # issue a replacement, then revoke the old one

create and rotate print the full secret exactly once, right after it’s active — save it immediately, it cannot be shown again:

Key created: work-laptop
  id:      a1b2c3d4e5f6
  created: 2026-07-23T16:32:10Z

Save this key now - it will not be shown again:

    sk-localai-9f1a2b3c4d5e6f7089...

Use it as a Bearer token:
  curl http://127.0.0.1:11435/v1/models \
    -H "Authorization: Bearer sk-localai-9f1a2b3c4d5e6f7089..."

Once at least one key is active, every request needs it:

curl http://127.0.0.1:11435/v1/models \
  -H "Authorization: Bearer sk-localai-REPLACE_ME"
from openai import OpenAI

client = OpenAI(
    base_url="http://127.0.0.1:11435/v1",
    api_key="sk-localai-REPLACE_ME",
)

Behavior notes:

API keys authenticate requests; they don’t encrypt them. For LAN/WAN access, put a TLS-terminating reverse proxy in front and restrict it with a firewall — see Security.

Documentation

Security

The helper scripts bind llama-swap to 127.0.0.1, so the API is available only on the local machine by default. Do not expose it to a network without adding authentication (localai key create), TLS, and appropriate firewall rules.

Credits

This project is built on top of:

Special thanks to the maintainers and contributors of these projects.

LocalAI focuses on simplifying installation, configuration, model management, and service deployment for local LLM environments.

Support

Buy me a coffee
If this repo helped you, give it a star