Qwen3.8-Flash-Next on an RTX 4090 and three RTX 3090s
Follow-up to the Qwen3.8-27B article. I experimented with Qwen3.8-Flash-Next, the MoE bigger brother, on two of the machines from last time: Ulmus (one RTX 4090, Ryzen 7900) and the triple RTX 3090 server at work. Flash-Next is a MoE (125B parameters), only 6B active per token, plus a 51B n-gram table and a small MTP head. So cheap compute but needs vast amount of RAM/VRAM. The plan was to keep all of it in RAM, and let the GPU hold the dense layers, the KV cache and a cache of…
Continue reading...