A take I've come to believe around generalization/sample efficiency is that if we do start to see AIs improving a lot at generalization/sample efficiency, we will see it first in robotics/robotics evals, and we have early signs of robotics/robotics evals starting to improve without robotics-specific data, which is bullish for generalization.
There are a couple of reasons for this:
1. A lot of the proposed use cases for robotics, like industry robots that handle varied environments or home robots require more generalization, especially from other data, unlike what has appeared to be the situation for tests and coding/easily grindable mathematics (and potentially white collar work.) Home robotics are an especially hard test case, because homes are spaces of infinite variation, long tails and constant change. As such, the scope of a home robot cannot be narrow. To create real value, it must perform reliably across open-ended variation with no special setup, and that is a practical test of embodied general intelligence.)
2. I'd say the major reason for this comes down to the fact that in areas where we don't need much generalization, we already have good enough robots for various tasks, and there's no overhang for AIs to take and become suddenly good without generalizing to other skills (unlike tests where IQ tests do measure human performance at lots of things somewhat well, or coding/easily grindable mathematics.)
3. One other thing that helps is that it's much easier to restrain LLM training data around the physical world, partially from the world itself being only partial information, but also the fact that it's much easier to restrain poorly generalizing AIs from just cheating it with a lot of training data, because it's easier to avoid training data leakage/easier to bound how much AIs need to generalize. It's also easier to give them hard problems, unlike benchmarks on the internet, because the model might just not be trained on robotics data (and this is fe