Download a pinned model below or choose your own compatible folder. Single checkpoints allow 4 GiB; experimental indexed shards allow 32 GiB total. Context stays capped at 512 tokens. CPU inference may be slow and generation has a five-minute limit for disk-backed models. GPU, GGUF, quantization and general 7B/14B certification are not included.
Conversation CRUD works here and is stored only in this browser. Model selection and generation are unavailable; no model or replies are simulated.
A workspace you own.
Model library — download, resume and load local models
Downloads require Internet and licence acceptance. Inference and conversations stay on your computer.
Start with a local conversation.
Create a conversation, bring a compatible model folder, then choose how your prompt should be handled. Your history stays here.
512 tokens