Models / qwen3-8b-fp8

Qwen3 8B FP8

Qwen3 8B with every decoder projection in FP8 E4M3, one scale per block of 128 by 128 weights: Qwen's own FP8 checkpoint. 60% of bf16's memory, faster decoding.

8.2B parametersqwen3Apache-2.0text-generationdecoder-onlygrouped-query-attentionqk-normfp8

No sample has been recorded for Qwen3 8B FP8 yet. What it produces, next to the reference stack's answer, will appear here once the benchmark harness has run it.