rnj-1-instruct08 Dec 2025
Ashish Vaswani's Essential, which has released a 24T token pre-training dataset in the open before, released an 8B Gemma-like model with strong agentic capabilities. The blog goes into more details why they decided to pre-train models and some of the technical decisions and findings.