From Multi-Agent Reinforcement Learning to Agentic AI: A Comprehensive Literature Review of Algorithmic Advances and Decision-Analytic Implications (2020-2025)
Published 2026-08-05
Keywords
- Multi-agent reinforcement learning,
- Agentic AI,
- Large language model agents,
- Cooperative MARL,
- Game theory
- Decision analytics,
- Autonomous decision-making,
- Centralized training decentralized execution,
- Communication learning,
- LLM-based multi-agent systems ...More
Copyright (c) 2026 Bharatendra Rai, Milena Popović (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.
How to Cite
Abstract
Agentic artificial intelligence has evolved from a research aspiration to a deployable technology between 2020 and 2025. This evolution rests on two intertwined research trajectories: the maturation of multi-agent reinforcement learning (MARL) for coordinated sequential decision-making, and the emergence of large language model (LLM)-based agent architectures integrating symbolic reasoning, tool use, and natural-language communication into cooperative multi-agent workflows. This literature review synthesizes 57 peer-reviewed and openly archived contributions published since 2019 across journals and reputable venues, organized into a thematic taxonomy spanning value-decomposition algorithms (QMIX, QPLEX, Weighted QMIX, FACMAC), trust-region and sequence-model policy methods (MAPPO, HAPPO, MAT, HARL, UPDeT), communication and role learning (NDQ, I2C, ROMA, RODE), credit assignment (LICA, Difference Rewards Policy Gradients, DOP), game-theoretic equilibrium solvers (Pipeline PSRO, JPSRO, Online Double Oracle), open-ended and mixed-motive learning (Open-Ended Learning Team, CICERO, alliance dilemmas), and LLM-based agentic frameworks (AutoGen, MetaGPT, CAMEL, AgentVerse, ChatDev, Generative Agents, Voyager, ReAct, Reflexion, Tree of Thoughts). We compare benchmark and reproducibility infrastructure (PettingZoo, EPyMARL benchmarking, SMAC variants), examine application domains (autonomous driving, multi-agent pathfinding, software engineering, scientific discovery), and discuss implications for applied decision analytics, including human-in-the-loop arbitration, risk-bounded coordination, and verifiable autonomy. We close with an agenda of open problems including non-stationarity, credit assignment under partial observability, alignment and safety in deceptive agents, evaluation under distribution shift, and integrating symbolic reasoning with reinforcement-learned policies to guide the next phase of agentic AI research.
Downloads
References
- Rashid, T., Samvelyan, M., Schroeder de Witt, C., Farquhar, G., Foerster, J., & Whiteson, S. (2020). Monotonic value function factorisation for deep multi-agent reinforcement learning. Journal of Machine Learning Research, 21(178), 1-51. arXiv:2003.08839.
- Mahajan, A., Rashid, T., Samvelyan, M., & Whiteson, S. (2019). MAVEN: Multi-agent variational exploration. Advances in Neural Information Processing Systems (NeurIPS), 32. arXiv:1910.07483.
- Zhang, K., Yang, Z., & Basar, T. (2021). Multi-agent reinforcement learning: A selective overview of theories and algorithms. In K. G. Vamvoudakis, Y. Wan, F. L. Lewis, & D. Cansever (Eds.), Handbook of reinforcement learning and control (pp. 321-384). Springer. arXiv:1911.10635.
- Oroojlooy, A., & Hajinezhad, D. (2023). A review of cooperative multi-agent deep reinforcement learning. Applied Intelligence, 53(11), 13677-13722. arXiv:1908.03963.
- Hughes, E., Anthony, T. W., Eccles, T., Leibo, J. Z., Balduzzi, D., & Bachrach, Y. (2020). Learning to resolve alliance dilemmas in many-player zero-sum games. Proceedings of the 19th International Conference on Autonomous Agents and MultiAgent Systems (AAMAS). arXiv:2003.00799.
- Marris, L., Muller, P., Lanctot, M., Tuyls, K., & Graepel, T. (2021). Multi-agent training beyond zero-sum with correlated equilibrium meta-solvers. Proceedings of the International Conference on Machine Learning (ICML). arXiv:2106.09435.
- Bakhtin, A., Brown, N., Dinan, E., Farina, G., Flaherty, C., Fried, D., Goff, A., Gray, J., Hu, H., Jacob, A. P., Komeili, M., Konath, K., Kwa, M., Lewis, M., Mishra, A., Renduchintala, A., Roller, S., Rowe, D., Shi, W., Spisak, J., Wei, A., Wu, D., Zhang, H., & Zijlstra, M. (2022). Mastering the game of no-press Diplomacy via human-regularized reinforcement learning and planning. Science, 378(6624), 1067-1074. arXiv:2210.05492.
- Rashid, T., Farquhar, G., Peng, B., & Whiteson, S. (2020). Weighted QMIX: Expanding monotonic value function factorisation for deep multi-agent reinforcement learning. Advances in Neural Information Processing Systems (NeurIPS), 33. arXiv:2006.10800.
- Wang, J., Ren, Z., Liu, T., Yu, Y., & Zhang, C. (2021). QPLEX: Duplex dueling multi-agent Q-learning. Proceedings of the International Conference on Learning Representations (ICLR). arXiv:2008.01062.
- Wang, T., Wang, J., Zheng, C., & Zhang, C. (2020). Learning nearly decomposable value functions via communication minimization. Proceedings of the International Conference on Learning Representations (ICLR). arXiv:1910.05366.
- Yu, C., Velu, A., Vinitsky, E., Gao, J., Wang, Y., Bayen, A., & Wu, Y. (2022). The surprising effectiveness of PPO in cooperative multi-agent games. Advances in Neural Information Processing Systems (NeurIPS), 35. arXiv:2103.01955.
- Kuba, J. G., Chen, R., Wen, M., Wen, Y., Sun, F., Yang, Y., & Wang, J. (2022). Trust region policy optimisation in multi-agent reinforcement learning. Proceedings of the International Conference on Learning Representations (ICLR). arXiv:2109.11251.
- Zhong, Y., Kuba, J. G., Feng, X., Hu, S., Ji, J., & Yang, Y. (2024). Heterogeneous-agent reinforcement learning. Journal of Machine Learning Research, 25, 1-67. arXiv:2304.09870.
- Peng, B., Rashid, T., Schroeder de Witt, C. A., Kamienny, P.-A., Torr, P. H. S., Böhmer, W., & Whiteson, S. (2021). FACMAC: Factored multi-agent centralised policy gradients. Advances in Neural Information Processing Systems (NeurIPS), 34. arXiv:2003.06709.
- Wang, Y., Han, B., Wang, T., Dong, H., & Zhang, C. (2021). DOP: Off-policy multi-agent decomposed policy gradients. Proceedings of the International Conference on Learning Representations (ICLR). arXiv:2007.12322.
- Zhou, M., Liu, Z., Sui, P., Li, Y., & Chung, Y. Y. (2020). Learning implicit credit assignment for cooperative multi-agent reinforcement learning. Advances in Neural Information Processing Systems (NeurIPS), 33. arXiv:2007.02529.
- Wen, M., Kuba, J. G., Lin, R., Zhang, W., Wen, Y., Wang, J., & Yang, Y. (2022). Multi-agent reinforcement learning is a sequence modeling problem. Advances in Neural Information Processing Systems (NeurIPS), 35. arXiv:2205.14953.
- Hu, S., Zhu, F., Chang, X., & Liang, X. (2021). UPDeT: Universal multi-agent reinforcement learning via policy decoupling with transformers. Proceedings of the International Conference on Learning Representations (ICLR). arXiv:2101.08001.
- Ding, Z., Huang, T., & Lu, Z. (2020). Learning individually inferred communication for multi-agent cooperation. Advances in Neural Information Processing Systems (NeurIPS), 33. arXiv:2006.06455.
- Zhu, C., Dastani, M., & Wang, S. (2024). A survey of multi-agent deep reinforcement learning with communication. Autonomous Agents and Multi-Agent Systems, 38, 4. arXiv:2203.08975.
- Wang, T., Dong, H., Lesser, V., & Zhang, C. (2020). ROMA: Multi-agent reinforcement learning with emergent roles. Proceedings of the International Conference on Machine Learning (ICML). arXiv:2003.08039.
- Wang, T., Gupta, T., Mahajan, A., Peng, B., Whiteson, S., & Zhang, C. (2021). RODE: Learning roles to decompose multi-agent tasks. Proceedings of the International Conference on Learning Representations (ICLR). arXiv:2010.01523.
- Castellini, J., Devlin, S., Oliehoek, F. A., & Savani, R. (2021). Difference rewards policy gradients. Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems (AAMAS). arXiv:2012.11258.
- Papoudakis, G., Christianos, F., Schäfer, L., & Albrecht, S. V. (2021). Benchmarking multi-agent deep reinforcement learning algorithms in cooperative tasks. NeurIPS Datasets and Benchmarks Track. arXiv:2006.07869.
- McAleer, S., Lanier, J., Fox, R., & Baldi, P. (2020). Pipeline PSRO: A scalable approach for finding approximate Nash equilibria in large games. Advances in Neural Information Processing Systems (NeurIPS), 33. arXiv:2006.08555.
- Dinh, L. C., Yang, Y., McAleer, S., Perez-Nieves, N., Slumbers, O., Tian, Z., Mguni, D. H., Bou Ammar, H., & Wang, J. (2022). Online double oracle. Transactions on Machine Learning Research (TMLR). arXiv:2103.07780.
- Yang, Y., Ma, C., Ding, Z., McAleer, S., Jin, C., Wang, J., & Sandholm, T. (2025). Game-theoretic multiagent reinforcement learning. Foundations and Trends in Machine Learning. arXiv:2011.00583.
- Stooke, A., Mahajan, A., Barros, C., Deck, C., Bauer, J., Sygnowski, J., Trebacz, M., Jaderberg, M., Mathieu, M., McAleese, N., Bradley-Schmieg, N., Wong, N., Porcel, N., Raileanu, R., Hughes-Fitt, S., Dalibard, V., & Czarnecki, W. M. (Open-Ended Learning Team). (2021). Open-ended learning leads to generally capable agents. arXiv:2107.12808.
- Terry, J. K., Black, B., Grammel, N., Jayakumar, M., Hari, A., Sullivan, R., Santos, L. S., Dieffendahl, C., Horsch, C., Perez-Vicente, R., Williams, N., Lokesh, Y., & Ravi, P. (2021). PettingZoo: A standard API for multi-agent reinforcement learning. NeurIPS Datasets and Benchmarks Track. arXiv:2009.14471.
- Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. Proceedings of the International Conference on Learning Representations (ICLR). arXiv:2210.03629.
- Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems (NeurIPS), 36. arXiv:2303.11366.
- Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., & Narasimhan, K. (2023). Tree of Thoughts: Deliberate problem solving with large language models. Advances in Neural Information Processing Systems (NeurIPS), 36. arXiv:2305.10601.
- Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., & Scialom, T. (2023). Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems (NeurIPS), 36. arXiv:2302.04761.
- Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., Zhao, S., Hong, L., Tian, R., Xie, R., Zhou, J., Gerstein, M., Li, D., Liu, Z., & Sun, M. (2024). ToolLLM: Facilitating large language models to master 16000+ real-world APIs. Proceedings of the International Conference on Learning Representations (ICLR). arXiv:2307.16789.
- Xu, Q., Hong, F., Li, B., Hu, C., Chen, Z., & Zhang, J. (2023). On the tool manipulation capability of open-source large language models. arXiv:2305.16504.
- Wu, Q., Bansal, G., Zhang, J., Wu, Y., Li, B., Zhu, E., Jiang, L., Zhang, X., Zhang, S., Liu, J., Awadallah, A. H., White, R. W., Burger, D., & Wang, C. (2023). AutoGen: Enabling next-gen LLM applications via multi-agent conversation. arXiv:2308.08155 (presented at COLM 2024).
- Li, G., Hammoud, H. A. A. K., Itani, H., Khizbullin, D., & Ghanem, B. (2023). CAMEL: Communicative agents for “mind” exploration of large language model society. Advances in Neural Information Processing Systems (NeurIPS), 36. arXiv:2303.17760.
- Hong, S., Zhuge, M., Chen, J., Zheng, X., Cheng, Y., Zhang, C., Wang, J., Wang, Z., Yau, S. K. S., Lin, Z., Zhou, L., Ran, C., Xiao, L., Wu, C., & Schmidhuber, J. (2024). MetaGPT: Meta programming for a multi-agent collaborative framework. Proceedings of the International Conference on Learning Representations (ICLR). arXiv:2308.00352.
- Qian, C., Liu, W., Liu, H., Chen, N., Dang, Y., Li, J., Yang, C., Chen, W., Su, Y., Cong, X., Xu, J., Li, D., Liu, Z., & Sun, M. (2024). ChatDev: Communicative agents for software development. Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL). arXiv:2307.07924.
- Chen, W., Su, Y., Zuo, J., Yang, C., Yuan, C., Chan, C.-M., Yu, H., Lu, Y., Hung, Y.-H., Qian, C., Qin, Y., Cong, X., Xie, R., Liu, Z., Sun, M., & Zhou, J. (2024). AgentVerse: Facilitating multi-agent collaboration and exploring emergent behaviors. Proceedings of the International Conference on Learning Representations (ICLR). arXiv:2308.10848.
- Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., & Anandkumar, A. (2023). Voyager: An open-ended embodied agent with large language models. Transactions on Machine Learning Research (TMLR). arXiv:2305.16291.
- Park, J. S., O’Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., & Bernstein, M. S. (2023). Generative agents: Interactive simulacra of human behavior. Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST). arXiv:2304.03442.
- Lin, B. Y., Fu, Y., Yang, K., Brahman, F., Bhagavatula, C., Ammanabrolu, P., & Choi, Y. (2023). SwiftSage: A generative agent with fast and slow thinking for complex interactive tasks. Advances in Neural Information Processing Systems (NeurIPS), 36. arXiv:2305.17390.
- Li, H., Chong, Y. Q., Stepputtis, S., Campbell, J., Hughes, D., Lewis, M., & Sycara, K. (2023). Theory of mind for multi-agent collaboration via large language models. Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP). arXiv:2310.10701.
- Talebirad, Y., & Nadiri, A. (2023). Multi-agent collaboration: Harnessing the power of intelligent LLM agents. arXiv:2306.03314.
- Du, Y., Li, S., Torralba, A., Tenenbaum, J. B., & Mordatch, I. (2024). Improving factuality and reasoning in language models through multiagent debate. Proceedings of the International Conference on Machine Learning (ICML). arXiv:2305.14325.
- Jiang, D., Ren, X., & Lin, B. Y. (2023). LLM-Blender: Ensembling large language models with pairwise ranking and generative fusion. Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL). arXiv:2306.02561.
- Zhang, J., Xu, X., Zhang, N., Liu, R., Hooi, B., & Deng, S. (2024). Exploring collaboration mechanisms for LLM agents: A social psychology view. Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL). arXiv:2310.02124.
- Hubinger, E., Denison, C., Mu, J., Lambert, M., Tong, M., MacDiarmid, M., Lanham, T., Ziegler, D. M., Maxwell, T., Cheng, N., Jermyn, A., Askell, A., Radhakrishnan, A., Anil, C., Duvenaud, D., et al. (2024). Sleeper agents: Training deceptive LLMs that persist through safety training. arXiv:2401.05566.
- Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., Zhao, W. X., Wei, Z., & Wen, J.-R. (2024). A survey on large language model based autonomous agents. Frontiers of Computer Science, 18(6), 186345. arXiv:2308.11432.
- Guo, T., Chen, X., Wang, Y., Chang, R., Pei, S., Chawla, N. V., Wiest, O., & Zhang, X. (2024). Large language model based multi-agents: A survey of progress and challenges. Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI). arXiv:2402.01680.
- Cheng, Y., Zhang, C., Zhang, Z., Meng, X., Hong, S., Li, W., Wang, Z., Wang, Z., Yin, F., Zhao, J., & He, X. (2024). Exploring large language model based intelligent agents: Definitions, methods, and prospects. arXiv:2401.03428.
- Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2024). SWE-bench: Can language models resolve real-world GitHub issues? Proceedings of the International Conference on Learning Representations (ICLR). arXiv:2310.06770.
- Kiran, B. R., Sobh, I., Talpaert, V., Mannion, P., Sallab, A. A. A., Yogamani, S., & Pérez, P. (2021). Deep reinforcement learning for autonomous driving: A survey. IEEE Transactions on Intelligent Transportation Systems, 23(6), 4909-4926. arXiv:2002.00444.
- Zhang, R., Hou, J., Walter, F., Gu, S., Guan, J., Röhrbein, F., Du, Y., Cai, P., Chen, G., & Knoll, A. (2024). Multi-agent reinforcement learning for autonomous driving: A survey. arXiv:2408.09675.
- Damani, M., Luo, Z., Wenbin, E., & Sartoretti, G. (2021). PRIMAL2: Pathfinding via reinforcement and imitation multi-agent learning - Lifelong. IEEE Robotics and Automation Letters, 6(2), 2666-2673. arXiv:2010.08184.
- Qu, G., Lin, Y., Wierman, A., & Li, N. (2020). Scalable multi-agent reinforcement learning for networked systems with average reward. Advances in Neural Information Processing Systems (NeurIPS), 33. arXiv:2006.06626.
