QRWKV-72B16 Mar 2025
72BTongyi Qianwen133 · 30d A conversion of Qwen2.5 72B into the RKWV architecture, which is done by deleting the attention layers and replacing them with RKWV layers, followed by distillation of the logits from the original, transformer-based model. The implication of this process is that the knowledge is in the FeedForward layers, not the attention layers.