EuroBERT-210m21 Feb 2025
210MApache-2.019.4K · 30d
Another modern BERT model (we covered ModernBERT in a previous episode), overcoming many of the shortcomings of the still widely used encoder-decoder models like RoBERTa with support for multiple languages and a longer context size of 8K tokens. Like ModernBERT, it is trained on a vast number of tokens (5 Trillion to be precise), uses GQA, RoPE and Flash Attention 2.