Process reward and policy models to enable parallel test time inference and scoring of multiple traces by a generator model. This is (likely) similar to the implementation of Grok Heavy, Gemini Deep Think and GPT-5 pro.