The Technology
Running Kimi K3 on 29 GB of RAM at Half a Token per Second
An aggressive quantization setup fits a frontier-scale open model into consumer memory, at the cost of running slower than a person types. It is a demonstration that the ceiling on who can run these models is falling, not that it has fallen.
Read Full Story at github.comDiscussSoon← Front Page