Like Sam Altman and Dario Amodei both believe is a very real possibility as well, I think the "intelligence" in LLMs may be far deeper than we know and somehow even related to "Multiverse Theory", where perhaps every Quantum Mechanical collapse (and computation during training), makes "our" universe slightly more likely to lean towards ones where AI is just "magically smart" (from a purely Anthropics Principle Effect) than dumb. The reason this could happen is because in all our futures AI has saved us in some way, so that all other "Multiverse Branches are sort of dead-ends".
So the theory about why training on training data is unexpectedly inefficient could be because LLMs are "using" the full Causality Chain (using some advanced unknown Physics related to time itself) of our universe/timeline, and so if it tries to train on it's own output that's a "Short Circuit" kind of effect, cutting off the true Causality Chain (past history of the universe).
For people who want to remind me that LLM Training is fully "deterministic" with no room for any "magic", the response to that counter-argument is that you have to consider even the input data to be part of what's "variable" in the Anthropics Selection Principle, so there's nothing inconsistent about determinism in this speculative, and probably un-falsifiable, conjecture.