web-llm-chat
Run large language models locally in the browser for private AI conversations.
- Stars
- 1,063
- Forks
- 225
- Updated
- Updated Jul 13, 2026
Run large language models locally in the browser for private AI conversations.
WebLLM Chat loads and runs language models inside a WebGPU-capable browser, keeping prompts, responses, and model processing on the user’s device instead of a cloud inference server. After the initial setup and model download, it can continue offline, with Markdown rendering, image-based conversations, and a selection of built-in models. Advanced users can also connect the interface to custom models served through MLC-LLM. Browser-native inference requires WebGPU support.
Resource types
Use cases
Runtime
Organize live and asynchronous team chat with topic threads
Use Matrix for messaging and collaboration through a web or desktop client.
Protocols & integrations
Audience
Public GitHub facts last synced Jul 14, 2026.
Deploy a private Llama 2 chat interface locally with Docker or umbrelOS.
Use Matrix messaging through a clean web client you can self-host.