RLP: Reinforcement as a Pretraining Objective https://research.nvidia.com/labs/adlr/RLP/ https://github.com/NVlabs/RLP/tree/main