Skip to main content

Documentation

Local-first LLM inference for your apps.

Using an AI coding agent? Give it the NobodyWho skill:

npx skills add https://github.com/nobodywho-ooo/nobodywho --skill nobodywho

Run open-weight language models directly inside your software. Streaming chat, tool calling, structured output, embeddings, text to speech, speech to text and RAG. All offline with GPU acceleration. No servers, no API keys, no Docker. Built on llama.cpp.

Get started.

New to local LLMs

Start here if you are new to running language models locally. These guides cover the core concepts — what models are, how to pick one, and how quantization works.

Choose your binding.

Batteries-included bindings with sync and async APIs.

pip install nobodywho
Read the Python docs