All models

OLMo-1B-0724-hf

byAi2Ai2· 15 Jun 2024
General purpose

We updated our OLMo 7B model with some small improvements to training data and annealing techniques. A pretty minor bump, but we have some exciting new models coming soon. An interesting thing I've learned about pretraining is how loss spikes often relate to "skipped tokens," making the models worse at a fixed compute budget. This is because the gradients is loss spikes get clipped and the data effectively does nothing.

Similarity · VAIL
VAIL
VAIL Fingerprint
03e7:05e2:071f:0961:0c51:10ba:1fc9:6422

Explore other models with behavioral similarity to OLMo-1B-0724-hf.

Adoption · Hugging Face
RAM score
Hugging Face Downloads
7.2K
last 30d
944.4K
all time
HF Likes
24

Relative Adoption Metric contextualizes downloads against the model's size bucket.

Related Models