The Technology
An 80B Qwen Model Runs in 4.3 GB of RAM on a Mac
A new quantization and streaming approach lets an 80-billion-parameter Qwen model run in roughly 4.3 GB of RAM on a Mac, with a 35B variant running on an iPhone. The demonstration continues a steady collapse in the hardware floor for large models, pushing capable inference from data centers toward devices people already own.
Read Full Story at github.comDiscussSoon← Front Page