All models

Nemotron-H-56B-Base-8K

byNVIDIANVIDIA· 14 Apr 2025
General purpose

A hybrid transformer-mamba model trained on a whopping 20T tokens. The report goes into greater detail. The models perform on the same level as an attention-only variant, while being more than twice as fast. The long-context benchmarks also look competitive, making this architecture a serious contender to replace models which use sliding window attention or similar variants. The released models, however, support only 8K context.

Specs
Params56B
LicenseNVIDIA Internal Scientific Research and Development Model License
Similarity · VAIL
VAIL
VAIL Fingerprint
0024:0047:0070:0094:00e9:0133:0364:408e

Explore other models with behavioral similarity to Nemotron-H-56B-Base-8K.

Adoption · Hugging Face
RAM @ 30d
Relative Adoption Metric: 0.00×. Below benchmark. Measured at 30 days.
Hugging Face Downloads
15.1K
last 30d
175.1K
all time
HF Likes
33

Relative Adoption Metric contextualizes downloads against the model's size bucket.

Related Models