Qwen3 8B FP8 has not been measured yet. Its numbers will appear here once the benchmark harness has run it on a GPU against the stacks people already use.
Models / qwen3-8b-fp8
Qwen3 8B FP8
Qwen3 8B with every decoder projection in FP8 E4M3, one scale per block of 128 by 128 weights: Qwen's own FP8 checkpoint. 60% of bf16's memory, faster decoding.
8.2B parametersqwen3Apache-2.0text-generationdecoder-onlygrouped-query-attentionqk-normfp8