Qwen3.8-Flash-Next on an RTX 4090 and three RTX 3090s

Follow-up to the Qwen3.8-27B article. I experimented with Qwen3.8-Flash-Next, the MoE bigger brother, on two of the machines from last time: Ulmus (one RTX 4090, Ryzen 7900) and the triple RTX 3090 server at work.

Flash-Next is a MoE (125B parameters), only 6B active per token, plus a 51B n-gram table and a small MTP head. So cheap compute but needs vast amount of RAM/VRAM.

The plan was to keep all of it in RAM, and let the GPU hold the dense layers, the KV cache and a cache of the most used experts. llama.cpp gave me ~24 tok/s on a Q4 file with the 4090. Strata produced 60 on the same file. With the IQ3_S quant I reached ~110. A very respectable value.

Using the 27b, Q4 was the floor. This one behaves differently. It's a different algo, and the published scores are the same as the unquantized model (LiveCodeBench 87 vs 87, GPQA Diamond 92 vs 93). My own checks converge: 28 to 30 out of 30 on my DevOps/Python tasks depending on the seed, and 14/15 on a set of images. Q4 was not better beyond the noise, twice slower, and doesn't even fit in the RAM of the triple 3090.

MeasurementRTX 40903x RTX 30902x RTX 3090
(27B)
Memory24 GB VRAM GDDR6X
192 GB DDR5
3 × 24 GB VRAM GDDR6X
96 GB DDR4
2 × 24 GB VRAM GDDR6X
Power cap280 W225 W per card (x3)225 W per card (x2)
RuntimeStrata / CUDAStrata / CUDAvLLM / CUDA
Quant / KVIQ3_S
int8 KV
IQ3_S
int8 KV
W4A16
BF16 KV
SpeculationMTP4MTP4DFlash2 k=7
Service context262,144262,144262,144
Decode @32K120 tok/s95 tok/s167 tok/s
(short prompt, 200 W per card)
Decode @128K112 tok/s84 tok/s 
Decode @256Knot measured78 tok/s 
~32K prefill~4,700 tok/s
(7 s to first token)
~2,550 tok/s
(12 s to first token)
~1,500 tok/s
(21 s to first token)
Power (inference)~240 W~365 W 
Energy per token~2 J~4 J~2.4 J
(at the 400 W cap)
27B for reference
(decode / prefill)
179 tok/s / 2,300 tok/s  

Not a strict comparison, as the model, runtime and quant are not the same as the 27B ones but still an interesting reference point.

One RTX 4090 (Ulmus)

~120 tok/s at 32K, and still 112 at 128K. Prefill is twice the 27B's, so a 32K prompt starts answering after 7 seconds instead of 14. Decode on the other hand is a third slower than the 27B with DFlash2.

The full 256K context works, and so does vision. The vision tower lends its VRAM to the expert cache between images. That's ~5% more decode speed for ~100 ms more on each new image. Total is 23.5 GB of VRAM and ~80 GB of RAM, at about 240 W on a 280 W cap. So ~2 J per token, where the 27B was at 1.6.

Everything is in the RTX 4090 Flash-Next repository.

Hardware: B650 + Ryzen 7900, 192GB DDR5.

Three RTX 3090 (work)

I thought maybethat this one would beat the 4090. 72 GB of VRAM holds 99% of the experts, so no more RAM or PCIe bottleneck. I deleted 152 GB of old models (GLM-4.5-Air, Qwen3.5-122B, Qwen3-Next-80B) to make room, and tried it out.

Experts do fit. It's still slower. The best layout is a split of the layers over two cards, with the third one (with only 8 lanes) only running vision. That gives ~95 tok/s at 32K and still 78 at 256K, with prefill at 2,500 to 3,200 tok/s. So 22 to 26% slower decode than the 4090, cold prompts 1.4 to 1.8x longer, and almost twice the energy per token.

Splitting layers doesn't speed up a single request, the token still goes through the cards one after the other. And each layer is slower on a 3090 anyways. I also tried :

  • One card + experts on the CPU : 73 tok/s, 23% slower than two cards.
  • Three-way split : same decode, 25% slower prefill, +90 W.
  • Q4 : doesn't fit in the 96 GB of RAM.

The cards stay at 225 W, as I've said multiple times, anything above that is factory overclocking. Quality is the same as on the 4090 (29/30 and 14/15). Flash-Next is now the default model of this machine and the 27B starts on demand, as they both need the same cards. Everything is in the triple RTX 3090 Flash-Next repository.

Hardware: Intel X99 + 6900k, 96GB DDR4.

Versus the 27B

Flash-Next is moderately smarter than the 27B. Our internal benchmark at work (75 questions) doesn't show much as it's reasonning around provided context : 60 vs 58 with reasoning on low, where GPT-6.1 Sol got 72 and Haiku 4.5 got 29. With reasoning off, both fall to ~35, so don't unless your need the latency (which I sometime do).

Where 3.8-Flash-Next shines is, without much surprise: knowledge.

Larger models hold more information. And depending on the type of task at hand, it can help. Typically when you don't or cannot provide all the context or vocabulary necessary to handle or understand a task properly. Or when the agent is not allowed internet search to fetch more (or accurate) information to accomplish its task. 

It's very interesting to see that "reasonning" has really been shoehorned into a 27B dense model with a lot a success and that it doesn't "need" more.

I've done dead reckoning regarding gpt-5.6-sol's size a few months ago, when they released it on Cerebras's infrastructure and guessed that it's likely in the range of 50 to 60B active parameter. Meaning that absolute cutting edge "AI" needed only twice what Qwen 3.8-27B has.
Of course, 5.6-sol is huge, and cannot be directly compared in knowledge. But this gives an interesting comparison point. Again, it was dead reckoning around a blackbox, don't take this at face value. 

The fairly recent advent of Fable and Astra can also be ballparked, by looking at their relative pricing to previous SoTA (respectively x5 between 6.1-sol and 6-astra, and 2.5x between opus-5.5 and fable-5.1). 

I must say that overall I'm a bit disappointed of not gaining more by uping from 27b to Flash-Next. I'll see if it fits on my Strix Halo, where it could really do wonders (as a dense model are impossible to run on it). After all, that's where MoEs feel at home.

Disclaimer: table been llm-written/assisted