Models that carry their architecture.
Every model in Nest is a checked .linnet source with a SafeTensors checkpoint on the Hugging Face Hub. Shapes and dtypes are verified before anything runs, and the same file loads in PyTorch, JAX, XLA, and ONNX Runtime.
from linnet import nest
model = nest.load("tinyllama-1.1b-chat", backend="torch")all-MiniLM-L6-v2
22.6MThe most downloaded model on the Hub: a 6-layer, width-384 BERT encoder distilled for sentence embeddings, mean-pooled and L2-normalized.
BERT Base (uncased)
109.5MGoogle's original BERT: 12 post-norm encoder layers of width 768, word/position/token-type embeddings, and the next-sentence-prediction pooler.
DINOv2-Base
86.6MMeta's self-supervised Vision Transformer: 14-pixel patches of a 518-pixel image, 12 layers of width 768 with LayerScale, trained without labels for retrieval, segmentation and depth features.
GPT-2 (124M)
124.4MOpenAI's 124M-parameter GPT-2: learned positions, pre-norm blocks, a fused QKV projection, and a head tied to the token embedding.
GPT-OSS 20B
20.9BA 21B-parameter mixture-of-experts decoder, 32 experts per layer with 4 active per token, alternating 128-position sliding-window and full causal attention with learned attention sinks, YaRN rope scaling for a 131072-token context, and expert weights read straight from the checkpoint's MXFP4 blocks and scales.
Llama 3.1 8B Instruct
8BAn 8B-parameter Llama-architecture decoder with grouped-query attention, llama3 rope scaling for a 131072-token context, a 128256-token vocabulary, and an untied output head, instruction-tuned.
Mistral 7B Instruct v0.3
7.2BA 7B-parameter Llama-architecture decoder with grouped-query attention, a 32768-token vocabulary and context, and an untied output head, instruction-tuned.
ModernBERT-base
149MAnswer.AI's 2024 redesign of the BERT encoder: 22 pre-norm layers of width 768, rotary positions, GeGLU, no biases, and a sliding attention window that opens up every third layer.
Phi-3 Mini 4K Instruct
3.8BA 3.8B-parameter decoder with fused QKV and gate/up projections, full (non-grouped) multi-head attention, RMS normalization, SwiGLU, and an untied output head, instruction-tuned for a 4K context.
Qwen2.5 0.5B Instruct
494.5MA 0.5B-parameter Qwen2 decoder with grouped-query attention, biased query/key/value projections, and an output head tied to the token embedding, instruction-tuned.
Qwen3 4B
4BA 4B-parameter Qwen3 decoder with grouped-query attention, per-head RMS normalization of the query and key projections, unbiased projections, and an output head tied to the token embedding.
Qwen3 8B
8.2BAn 8B-parameter Qwen3 decoder with grouped-query attention, per-head RMS normalization of the query and key projections, unbiased projections, and an untied output head.
ResNet-18
11.7MThe 18-layer residual network for ImageNet-1k: a 7x7 stem, four stages of two residual blocks, and a linear classifier.
ResNet-50
25.6MThe 50-layer residual network (v1.5) for ImageNet-1k: a 7x7 stem, four stages of bottleneck blocks 3, 4, 6, and 3 deep, and a linear classifier.
RoBERTa Base
124.1MMeta's more thoroughly trained BERT: the same post-norm encoder, dynamic masking, no next-sentence-prediction, and a byte-level BPE vocabulary of 50265.
SAM ViT-Base
93.7MMeta's Segment Anything Model: a ViT-Det image encoder with windowed and relative-position attention, whose 64x64 embedding any number of point or box prompts are decoded against by a two-way transformer.
SD VAE ft-MSE (decoder)
49.5MStability AI's MSE-finetuned autoencoder for Stable Diffusion, decoder half: group-normalized residual blocks, one spatial self-attention, and nearest-neighbour upsampling turn a 4-channel latent into an RGB image.
SDXL Base UNet
2.6BThe 2.6B-parameter denoising UNet of Stable Diffusion XL base: residual blocks and spatial transformers at three resolutions, cross-attending to a 2048-wide text embedding, conditioned on the timestep and on SDXL's pooled-text-plus-micro-conditioning vector.
SigLIP Base/16 224
203.2MGoogle's image-text model that replaces CLIP's softmax contrastive loss with a sigmoid one: 16-pixel patches through a headless ViT tower and a 64-token bidirectional text tower, meeting only in a scaled, biased dot product.
SmolLM2 1.7B Instruct
1.7BA 1.7B-parameter Llama-architecture model trained on 11 trillion tokens, instruction-tuned, with the head tied to the embedding.
TinyLlama 1.1B Chat v1.0
1.1BA 1.1B-parameter Llama 2 architecture trained on 3 trillion tokens, chat-tuned.
ViT-Base/16 224
86.6MGoogle's Vision Transformer for ImageNet-1k classification: 16-pixel patches of a 224-pixel image, 12 layers of width 768.
Whisper large-v3
1.5BOpenAI's largest multilingual speech recognition model: a convolutional stem over a 128-bin log-mel spectrogram, 32 encoder layers, and a 32-layer decoder that cross-attends to them, at width 1280.
Whisper tiny
37.8MOpenAI's smallest multilingual speech recognition model: a convolutional stem over a log-mel spectrogram, four encoder layers, and a four-layer decoder that cross-attends to them.