Models / modernbert-base

ModernBERT-base

Answer.AI's 2024 redesign of the BERT encoder: 22 pre-norm layers of width 768, rotary positions, GeGLU, no biases, and a sliding attention window that opens up every third layer.

149M parametersmodernbertApache-2.0fill-maskfeature-extractionencoder-onlyrotary-embeddingssliding-window-attention

Cosine similarity between 6 sentences. Not a trained sentence embedder: the last hidden state, mean-pooled over tokens and L2-normalized by hand, since this checkpoint has no embedding head.

  1. The cat curled up on the warm windowsill and fell asleep in the sun.
  2. Our new kitten refuses to eat anything except wet food.
  3. The central bank raised interest rates again to cool inflation.
  4. She rebalanced her portfolio after the market's sharp swings.
  5. The function threw an exception when the array index ran out of bounds.
  6. He refactored the module to remove three copies of the same logic.
123456
11.000.920.880.910.890.91
20.921.000.880.890.890.89
30.880.881.000.930.870.90
40.910.890.931.000.880.92
50.890.890.870.881.000.92
60.910.890.900.920.921.00

Shading runs from 0 to 0.93, the closest pair of different sentences.