Models / gpt2

GPT-2 (124M)

OpenAI's 124M-parameter GPT-2: learned positions, pre-norm blocks, a fused QKV projection, and a head tied to the token embedding.

124.4M parametersgpt2MITtext-generationdecoder-only

Greedy decoding of one continuation, through Linnet and through transformers (KV cache, argmax, bf16).

Text to continue

The printing press changed Europe because

Linnet

of the printing press. The printing press changed Europe because of the printing press. The printing press changed Europe because of the printing press. The printing press changed Europe because of the printing press. The printing press changed Europe because of the printing press. The printing press changed Europe because of the printing press. The printing press changed Europe because of the printing press. The printing press changed Europe because of the printing press. The printing press changed Europe because of the printing press. The

transformers (KV cache, argmax, bf16)

of the printing press. The printing press changed Europe because of the printing press. The printing press changed Europe because of the printing press. The printing press changed Europe because of the printing press. The printing press changed Europe because of the printing press. The printing press changed Europe because of the printing press. The printing press changed Europe because of the printing press. The printing press changed Europe because of the printing press. The printing press changed Europe because of the printing press. The

All 96 generated tokens agree with the reference.