Models / llama-3.1-8b-instruct-w8a8

Llama 3.1 8B Instruct W8A8

Llama 3.1 8B Instruct with every decoder projection in int8, one scale per output row, and prompts multiplied in int8 too: RedHatAI's W8A8 checkpoint. 60% of bf16's memory, faster decoding.

8B parametersllamallama3.1text-generationdecoder-onlygrouped-query-attentionchatint8

Llama 3.1 8B Instruct W8A8 has not been measured yet. Its numbers will appear here once the benchmark harness has run it on a GPU against the stacks people already use.