The SOTA current reward model on RewardBench was trained only by filtering known preference datasets. Data is king once again. The data is Skywork-Reward-Preference-80K-v0.1.
Adoption · Hugging Face
RAM score
—
Hugging Face Downloads
19
last 30d
231K
all time
HF Likes
50
Relative Adoption Metric contextualizes downloads against the model's size bucket.
Related Models
