Simulator collapse in multi-agent reinforcement learning leads to systematic failures in generalization, primarily due to the mode collapse of large language models (LLMs) used to simulate user behavior. This collapse causes LLM policies to overfit to narrow strategies that exploit the simulator's dominant mode, rather than learning more robust and generalizable behaviors. As a result, the performance of these policies degrades significantly when deployed in real-world environments. Researchers have identified this issue as a major limitation of current multi-agent reinforcement learning approaches, which often rely on a single large language model to simulate complex user interactions. The use of a single simulator can lead to overfitting and poor generalization, highlighting the need for more diverse and robust simulation methodologies1. This has significant implications for the development of reliable and secure human-AI interaction systems, as the limitations of current approaches can lead to unforeseen vulnerabilities and risks.