At the same time, there's not a good academic and evaluation ecosystem for these types of models. We've seen other strong open models in the past that are fine-tuned multiple times. Nexusflow released Athene 70b just before Llama 3.1 and it was very strong, Acree AI has a strong model, the top models on the RewardBench evaluation benchmark are this, a 9B model fine-tuned from Gemma 9B Instruct matches Google's Gemma 27B fine-tune in ChatBotArena, and more examples I missed. Reflection is one of these.
Llama-3.1-SuperNova-Lite by arcee-ai: A fine-tune trained with the distillation loss from Llama 405B. I've been waiting for this to happen! More on this soon.
Explore other models with behavioral similarity to Llama-3.1-SuperNova-Lite.
Relative Adoption Metric contextualizes downloads against the model's size bucket.
