Stabilizing Reinforcement Learning with LLMs: Formulation and Practices https://arxiv.org/abs/2512.01374 https://www.alphaxiv.org/ru/overview/2512.01374