Run models locally

Inference, UIs, and the Hugging Face stack on your machine.

llama.cpp

Efficient C/C++ LLM inference — run models locally on CPU and consumer GPUs with quantization

repo

Ollama

Run LLMs locally with a single command — supports Llama, Mistral, Gemma, and dozens more

tool

Hugging Face

The leading platform for ML models, datasets, and demos — hub for the open-source AI community

tool

vLLM

High-throughput LLM serving engine with PagedAttention for efficient memory management

repo

LM Studio

Desktop app to run any GGUF model locally — one-click downloads, chat UI, and local server mode

chat

Jan

Open-source offline-first desktop chat — runs Llama, Mistral, and custom models completely on-device

chat

Open WebUI

Self-hosted Ollama and OpenAI-compatible chat UI — full featured, RAG, multimodal, and pipelines

chat