We updated our OLMo 7B model with some small improvements to training data and annealing techniques. A pretty minor bump, but we have some exciting new models coming soon. An interesting thing I've learned about pretraining is how loss spikes often relate to "skipped tokens," making the models worse at a fixed compute budget. This is because the gradients is loss spikes get clipped and the data effectively does nothing.
Similarity · VAIL
VAIL Fingerprint
03e7:05e2:071f:0961:0c51:10ba:1fc9:6422
Explore other models with behavioral similarity to OLMo-1B-0724-hf.
Adoption · Hugging Face
RAM score
—
Hugging Face Downloads
7.2K
last 30d
944.4K
all time
HF Likes
24
Relative Adoption Metric contextualizes downloads against the model's size bucket.
Related Models
