Rethinking On-Policy Distillation of Large Language Models II: One Training Example https://arxiv.org/abs/2609.04172 https://www.alphaxiv.org/ru/abs/2609.04172