Transitive RL: Value Learning via Divide and Conquer https://arxiv.org/abs/2510.22512 alphaxiv.org/overview/2510.22512v1 https://seohong.me/blog/rl-without-td-learning/