The study investigated model collapse dynamics within multi-model ecosystems. Thirteen open 1–4B models were organized into ecosystems of 3 to 13 players. An injected probe was used to artificially increase the share of one model to 90%. Each model’s output was mixed into a shared pool based on market share, and models were retrained for five generations. The research found no significant difference in the speed of collapse or the final destination of the models regardless of the market share distribution.
Specifically, a more unequal split between models within an ecosystem only resulted in minor shifts in the five-generation endpoint, with a shift of only a few percent of the drift common to all arms. An extreme share paired with a strong injected bias did not guarantee steering, and the resulting topic shift left only a faint trace on the collapse measurement.
The speed of collapse was primarily determined by the source of the text in the pool and the susceptibility of the members of the ecosystem. Swapping members of a K=3 ecosystem changed five-generation drift by 2.8x when a share was held fixed. A share-weighted index of each member’s susceptibility explained the speed differences across nineteen arms with R^2 = 0.68. Replacing half the pool with human text roughly halved drift without changing its course.
These findings suggest that the composition of the training pool, rather than the concentration of models, is the primary driver of model collapse. The research highlights the importance of understanding the sources and characteristics of data used in multi-model training environments.
Source: https://arxiv.org/abs/2609.11146