llama.cpp
Efficient C/C++ LLM inference — run models locally on CPU and consumer GPUs with quantization
repoInference, UIs, and the Hugging Face stack on your machine.
Efficient C/C++ LLM inference — run models locally on CPU and consumer GPUs with quantization
repoState-of-the-art ML library with thousands of pretrained models for text, vision, and audio
repoThe leading platform for ML models, datasets, and demos — hub for the open-source AI community
toolDesktop app to run any GGUF model locally — one-click downloads, chat UI, and local server mode
chatOpen-source offline-first desktop chat — runs Llama, Mistral, and custom models completely on-device
chatSelf-hosted Ollama and OpenAI-compatible chat UI — full featured, RAG, multimodal, and pipelines
chat