On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models https://arxiv.org/abs/2512.07783