Models / sam-vit-base

SAM ViT-Base

Meta's Segment Anything Model: a ViT-Det image encoder with windowed and relative-position attention, whose 64x64 embedding any number of point or box prompts are decoded against by a two-way transformer.

93.7M parameterssamApache-2.0image-segmentationvisionpromptedencoder-decoder
Input
Input
Output
Output

A point prompt at (490.0, 170.0) on the right cat: Linnet's mask (left, blue) vs transformers.SamModel's (right, orange). IoU between the two masks: 1.000; predicted IoU, Linnet 0.979, reference 0.979.