All models

EuroBERT-210m

byEuroBERTEuroBERT· 21 Feb 2025
Embedding

Another modern BERT model (we covered ModernBERT in a previous episode), overcoming many of the shortcomings of the still widely used encoder-decoder models like RoBERTa with support for multiple languages and a longer context size of 8K tokens. Like ModernBERT, it is trained on a vast number of tokens (5 Trillion to be precise), uses GQA, RoPE and Flash Attention 2.

Specs
Params210M
LicenseApache-2.0
Adoption · Hugging Face
RAM score
Relative Adoption Metric not applicable.
Hugging Face Downloads
19.4K
last 30d
253.2K
all time
HF Likes
87

Relative Adoption Metric not applicable.

Related Models