KEYO StudioLocal workspace
Connecting
Own-engine CPU alpha.

Download a pinned model below or choose your own compatible folder. Single checkpoints allow 4 GiB; experimental indexed shards allow 32 GiB total. Context stays capped at 512 tokens. CPU inference may be slow and generation has a five-minute limit for disk-backed models. GPU, GGUF, quantization and general 7B/14B certification are not included.

Local inference / inspect your run

A workspace you own.

Active weights No model loaded
Model library — download, resume and load local models

Downloads require Internet and licence acceptance. Inference and conversations stay on your computer.

No conversation selected

Start with a local conversation.

Create a conversation, bring a compatible model folder, then choose how your prompt should be handled. Your history stays here.

Create a conversation to begin. Ctrl + Enter to run
Maximum generated tokens, from 1 to the alpha limit of 128.
Sampling temperature from 0 to 2.
Context limit
512 tokens