Strata, built by GitHub user Niko1221, runs a 125-billion-parameter LLM on a consumer gaming PC. The project targets Qwen3.8-Flash-Next, a mixture-of-experts architecture with 24,576 small experts, activating only 10 per token. Minimum specs are 32 GB of system RAM, 12 GB of VRAM, and 80 GB of storage.

The memory strategy is the technical core worth understanding: VRAM acts as a cache for frequently used experts, the full model lives in system RAM, and a 29 GB lookup table sits on SSD and is paged in as needed. Speculative decoding via multi-token prediction lets the model check several candidate tokens per pass. On an RTX 5070 with 12 GB VRAM, throughput lands between 50 and 90 tokens per second depending on quantization.

Once running, Strata serves a browser UI and exposes OpenAI- and Anthropic-compatible APIs on localhost, making it a drop-in backend for existing coding assistants and chat frontends. Factual recall had notable gaps in testing, but code analysis and porting tasks held up. The full write-up details setup friction with NVIDIA's C compiler and gcc headers, which is reason enough to read before you run setup.sh.

[READ ORIGINAL →]