Nanite for Experts: a fidelity-first memory hierarchy for local mixture-of-experts inference
Our full engineering research report. We treat NVMe, RAM, pinned host memory and VRAM as one coordinated hierarchy so a 35B sparse model can run on a 10 GB card — then show, with measurements, why most of the aggressive caching ideas were slower than simply placing layers correctly. Includes the negative results.
Read the report