Пока ходил искал алгоритмы которые в студию завести полезно собрал список *Zero алгоритмов Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm https://arxiv.org/abs/1712.01815 https://www.alphaxiv.org/ru/abs/1712.01815 Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model https://arxiv.org/abs/1911.08265 https://www.alphaxiv.org/ru/abs/1911.08265 Online and Offline Reinforcement Learning by Planning with a Learned Model https://arxiv.org/abs/2104.06294 https://www.alphaxiv.org/ru/abs/2104.06294 Learning and Planning in Complex Action Spaces https://arxiv.org/abs/2104.06303 https://www.alphaxiv.org/ru/abs/2104.06303 Mastering Atari Games with Limited Data https://arxiv.org/abs/2111.00210 https://www.alphaxiv.org/ru/abs/2111.00210 EfficientZero V2: Mastering Discrete and Continuous Control with Limited Data https://arxiv.org/abs/2403.00564 https://www.alphaxiv.org/ru/abs/2403.00564 ReZero: Boosting MCTS-based Algorithms by Backward-view and Entire-buffer Reanalyze https://arxiv.org/abs/2404.16364 https://www.alphaxiv.org/ru/abs/2404.16364 UniZero: Generalized and Efficient Planning with Scalable Latent World Models https://arxiv.org/abs/2406.10667 https://www.alphaxiv.org/ru/abs/2406.10667 One Model for All Tasks: Leveraging Efficient World Models in Multi-Task Planning https://arxiv.org/abs/2509.07945 https://www.alphaxiv.org/ru/abs/2509.07945 Object-Centric World Models Meet Monte Carlo Tree Search https://arxiv.org/abs/2601.06604 https://www.alphaxiv.org/ru/abs/2601.06604 PriorZero: Bridging Language Priors and World Models for Decision Making https://arxiv.org/abs/2605.12289 https://www.alphaxiv.org/ru/abs/2605.12289 Ну и либа в которой большая часть из них сделана и ребята новые алгоритмы создают время от времени: https://github.com/opendilab/LightZero