
Qwen3.8-27B on four local inference machines
Here is my experience with Qwen3.8-27B, the first small model I have found to be actually useful for agentic tasks. It has good general knowledge, solid tool use, stays on track even after multiple context compactions and overall gets the job done.
I've seen many telltale that it's been trained on some version of Opus. Such as "Excellent research" after a research session and being proud of having bypassed a harness restriction.
I tested it on four very different machines: an Intel Arc Pro B70 at work (~2.5k€ machine), a dual RTX 3090 at work (~2.5k€ machine), Ulmus my AI desktop computer (~6k€ machine), and Juniperus, my Ryzen AI Max+ 395 laptop (~6k€ machine).

This is not a strict GPU comparison as quantization and runtimes differ. But it reflects the compromises I ended up with on each machine.
In case you're not familiar with LLM inference jargon:
- Prefill is how fast the model converts the prompt into a mathematical representation.
- Decode is how fast it outputs tokens.
Prefill is compute-bound important to process large inputs in a decent amount of time. Typically .md files, vision projections (images), or tool-call results. Decode is often memory bandwidth bound.
| Measurement | Arc Pro B70 | RTX 4090 | 2x RTX 3090 | Ryzen AI Max+ 395 |
| Memory | 32 GB | 24 GB | 48 GB | 128 GB |
| Die size | 368 mm² | 609 mm² | 2 × 628 mm² | 308 mm² |
| Memory bandwidth | 608 GB/s | 1,008 GB/s | 2 × 936 GB/s (local) | 256 GB/s (shared) |
| Power cap | 210 W | 280 W | 450 W (2x 225 W) | 45 W (measured) |
| Runtime | vLLM / XPU | vLLM / CUDA | vLLM / CUDA | LM Studio / llama.cpp Vulkan |
| Benchmark quant / KV | GPTQ INT4 W4A16 FP8 KV | W4A16 BF16 KV | W4A16 BF16 KV | UD-Q6_K_L F16 KV |
| Speculation | MTP4 | DFlash2 k=7 | DFlash2 k=7 | MTP4 |
| Service context | 262,144 | 262,144 | 262,144 | 262,144 |
| Decode | 85.61 tok/s | 178.7 tok/s | 167.2 tok/s (400 W cap) | 12.60 tok/s |
| Decode efficiency (tok/s/W) | 0.408 (at cap) | 0.641 (at cap) | 0.418 (400 W cap) | 0.280 (45 W) |
| ~32K prefill | 1,418 tok/s | 2,299 tok/s | ~1,400 tok/s | <182 tok/s TTFT >180 s |
| Prefill efficiency (tok/s/W) | 6.75 (at cap) | 8.27 (at cap) | ~3.5 (400 W cap) | <4.04 (45 W) |
Quantization
I also checked whether Q4 had actually hurt the model
| Controlled comparison | BF16 (unquantized) | 4-bit |
| GPQA Diamond | 89.9% | Q4_K_M: 90.4% |
| IFBench | 81.7% | Q4_K_M: 81.0% |
| Terminal-Bench 2.1 | 63/89 (70.8%) | Q4_K_M: 63/89 (70.8%) |
| MMLU-Pro sample | 81.9% | W4A16: 82.6% |
| HumanEval | 93.9% | W4A16: 95.7% |
| IFEval strict | 81.89% | NF4: 81.89% |
| BFCL tool calling | 82.18% | NF4: 81.13% |
Q4 preserves all capability of the unquantized model. Going down to Q3 I started to see the model degrade, so Q4 is clearly the floor and I wouldn't risk getting under.
Intel Arc Pro B70
The B70 gives a very sane balance: 85 tok/s, usable prefill and 254K context. Depending on how much you paid for it, it's surprisingly decent silicon. I paid mine 1400€ VAT inc. and kinda regret not getting one when they were released at a lesser price. It's still honest pricing, and has the advantage of being available for purchase, contrary to the RTX 5090 FE that I'd want. Note that I also considered the AMD AI Pro R9700, but for similar perf, I got better price and better cooling (Asrock Creator has a good vapor chamber).
Anyways, getting there was painful as llama.cpp with Vulkan ran at around 10 tok/s. The final setup, brute-forced by GPT-5.6-sol, uses a patched vLLM/XPU container, GPTQ INT4, MTP4 and FP8 KV. I found pretty much no quality compromise with this setup. The complete tuned stack is available in my Intel Arc Pro B70 inference repository.
This is now our default nightly job model and machine at work, sitting in an air conditionned server room. It replaces a 3x3090 setup that used to run Qwen-Next-80b and Qwen 3.6-122b-a10b, without loosing in quality on our workloads and in-house benchmarks.
Hardware: B550, 64GB DDR4, Ryzen 5500GT.
One RTX 4090 (Ulmus)
I initially had poor results again with llama.cpp. My self-contained Docker/vLLM stack, adapted from syv-ai's single-GPU work, now serves Huihui's abliterated W4A16 target with much better performance, reaching almost 180 tok/s decode and 2300 tok/s prefill.
The deployed profile keeps vision and prefix caching while exposing the full context window thru DFlash2 k=7 and KVarN K4V2 KV.
The 0.85 GiB vision tower stays in pinned system RAM, but image computation still runs on the GPU: only its weights travers on the PCIe once per image. It's a modet 50ms tag. The complete tuned stack is available in my RTX 4090 inference repository.
This is my "at home" go-to model. For example, I've got an agent capable of accepting "Download {whatever movie name}", reach for torrenting sites, identify the quality range I prefer, pull the magnet link, invoke transmission-cli, follow the download, rename the downloaded movie cleanly, and upload it to my fileserver. This is made possible by the abliterated version that doesn't care about IP. It is also a useful red-teaming model.
Hardware: B650 + Ryzen 7900, 192GB DDR5.
Two RTX 3090
This machine hosts different models at work and is (well, was...) used to handle async jobs at night. It mostly ran Qwen3-Next-80B and Qwen3.5-122B. But as it sits mostly unused during the day, I tested how Qwen3.8-27B would behave.
Turns out splitting the model across only the two RTX 3090s running at x16 and moving to vLLM was a much, much better choice. The final W4A16 Huihui + DFlash2 setup reaches almost 170 tok/s at 200 W/card, against 38 tok/s with llama.cpp on three cards (the third using 8 lanes). It also reaches 270 tok/s aggregate with two streams. The cards are capped at 225 W each for slightly lower latency. The complete tuned stack is available in my dual RTX 3090 inference repository.
This setup has now become experimental (superseeded by the B70 as a workhorse), but being able to handle up to 4 streams, it could become a coding machine if for some reason, major coding models become unavailable. It can also help diagnose infrastructure issues quickly when the issue causes an internet outage.
Hardware: Intel X99 + 6900k, 96GB DDR4.
Ryzen AI Max+ PRO 395 / Radeon 8060S (Juniperus)
Juniperus (an HP ZBook G1a) ran a Q6 model with MTP4. Decode reached ~12 tok/s. A 32K test prompt didn't output anything before my three-minute timeout. The APU hovered at 45 W throughout the test (but I'm pretty sure I'm having a PSU wattage detection issue that accidentally caps the PPT).
The 128 GB shared memory makes almost any model fit, but getting throughput is something else. Qwen3.8-27B is dense, so it hammers compute and memory bandwidth. Both are fairly constrained on Strix Halo.
It can work as a short Q&A model, but not as a coding agent with a growing context. Dropping the quant to Q4 or lower and exploring vLLM optimization would likely increase performances. But the gap to "agentic" token rate is too wide for me to waste time on this. Plus, for Q&A, I can load much larger models with fewer active parameters, like Qwen3.8-Flash-Next.
I'm watching Halogen, as progress here could make Qwen-3.8-Flash-Next viable for agentic work. Preliminary tests on my side reached about 1000 tok/s prefill and 40 tok/s decode.
Disclaimer: table and charts been llm-written/assisted.