this discussion I had with alex earlier this year re-shaped a lot how I’m seeing these new breed of frontier models that are specialist at coding + multi-agent + long running.
the concepts of compositional generalization and sub tasks being locally in-distribution makes it very very obvious why LLM + harness are doing so well right now.
and also explain why tweaking the harness a bit sometime make a whole set of tasks that were previously unatainable for LLM a breeze.
I strongly believe now that processing tasks into a “coding-friendly” format will be a continuous vector of breakthrough for the months to come.
check out the full interview below and big thanks to everyone that submitted questions was an awesome session!
table of content:
0:00 - RLM are compositional generalizers
4:40 - overview of what is an RLM
10:00 - updated info and experiments on RLM
22:40 - Big Boss Alex L. Zhang
27:45 - what is “locally in-distribution”?
34:00 - do we even need long context inside an LLM?
36:50 - yacine is rambling again…
39:35 - is locally in-distribution computation enough for scientific breakthrough
46:40 - why are transformers poor at compositional generalization
52:50 - what is missing in this soup of capability from the transformers side.
57:25 - [at gun point] improve the harness or train a new model for compositional generalization???
1:03:00 - how early RLM should enter the training pipeline?
1:07:00 - how do you validate such a major architectural change btw?
1:10:00 - what does Alex thinks about long-running persistent ai agents?
1:14:00 - is code enough for compositional generalization across modalities?
1:18:52 - how are we doing at the frontier for long horizon tasks?
1:22:00 - speculative programmatic tool call (sPTC)
1:28:00 - does the RLM impose an abstraction the model should learn for itself?
1:32:00 - what’s next for Big Boss Alex L. Zhang (and what’s not next)?









