A 14-year-old Dell PowerEdge R720 server with no GPU is running GLM 5.3 Flash, Qwen 3.8 Flash, and Qwen 3.27B. Builder MattMo documented the setup on video. The hardware costs roughly $600 on the second-hand market.
The ceiling is 4 tokens per second, pushed out by dual Xeon processors across 20 threads. The models live entirely in 348 GB of DDR3 system memory. The CPUs are old enough to lack modern instruction sets that would accelerate inference, and the thread count is the primary bottleneck. MattMo does not pretend this is fast. He argues the case for batch workloads where speed is not the constraint, and for homelab operators who already own hardware in this class.
The interesting argument in the full video is not that this is practical for most people. It is that the RAM capacity threshold for flagship models is now crossable without a single dollar spent on VRAM. Read the original for MattMo's specific model configuration details, his take on whether upcoming model optimizations can move the needle past 4tps, and the honest accounting of who this actually makes sense for.
[READ ORIGINAL →]