Project overview
WebLLM Chat loads and runs language models inside a WebGPU-capable browser, keeping prompts, responses, and model processing on the user’s device instead of a cloud inference server. After the initial setup and model download, it can continue offline, with Markdown rendering, image-based conversations, and a selection of built-in models. Advanced users can also connect the interface to custom models served through MLC-LLM. Browser-native inference requires WebGPU support.
Repository facts
- Primary language
- TypeScript
- License
- Apache-2.0
- Repository updated
- Jul 13, 2026
- Default branch
- main
Resource types
Web app
Use cases
Communication and collaboration
Runtime
BrowserDockerLocal
Protocols & integrations
REST
Audience
General users