Skip to content

FastFlowLM NPU results: Gemma 4 E4B ​

Numbers from 2026-09-16 on a GMKtec EVO-X2 (Ryzen AI MAX+ 395 / XDNA2), FastFlowLM 1.0.5, --pmode performance. Deploy steps: FastFlowLM deploy.

Official reference: Gemma 4 NPU benches (Kraken Point / 32 GB, FLM 0.9.40). Raw table: bench CSV.


Service smoke (8219) ​

Short prompt “reply with one word: ok” returned ok. Client latency 1.52 s, TTFT 1.35 s, RSS ~9.06 GiB. Vision probe (red circle + blue rectangle) was correct.


1k–32k ladder (flm bench) ​

ContextTTFT (s)Prefill (tok/s)Decode (tok/s)Official Kraken decodeOfficial Kraken prefill
1k2.176449.511.7512.6441
2k3.479559.711.3812.3572
4k6.175629.110.9411.6668
8k11.685664.09.9510.6720
16k24.293638.38.549.0695
32k56.764546.16.616.8586

Decode tracks official Kraken, about 3–7% slower. Short-chat decode (~12.5 tok/s) is empty-KV; use this ladder for long-context decode. Short-chat prefill is not a steady-state bandwidth number.

Concurrency 1 / 2 / 4 on short ok all succeeded (no 503); wall clock ~1× / 2.2× / 4.8× because NPU runs one request at a time and queues the rest.