0:00
/
Generate transcript
A transcript unlocks clips, previews, and editing.

Pretrained LLMs Are Surrounded by Task Experts feat Yulu Gan from MIT

discussion with the great Yulu Gan about the shape of pre-training weights distribution

sometime there is an idea that just makes other weird research results make so much more sense and sort of cascade to explain a whole bunch of phenomenon.

neural thicket from the great yulu gan is one such paper that I absolutely loved from start to finish.

the old mental model I had for pre-trained neural networked was that the pre-training regime made them land on a single good set of weight that was optimal for many tasks.

but after going through the paper "Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights" my view has been changed.

the better idea seems to be that the model after pre-training is in a basin where it is surrounded by more specialized expert at varying distance from that set of weights which the authors calls thicket.

a bit like the valley surrounded by mountains.

now depending on whether the model has big or small capacity this density of thicket will be varying and these expert weights will also have varying level of difficulty to be reached via post-training.

this finding elegantly presented by Yulu & co. has far reaching consequence on how we should think about pre-training and it's connection with post-training models (and even distillation / continual learning)

very densely packed interview in there and there is a few open threads that are ripe to be picked up if you are interested in this research direction!

ps: my favorite bit is the connection with the baldwin effect in biology which is an evolutionary concept proposing that learned behaviors can shape the course of evolution, very fitting with post-training dynamics!

important links:
👉 yulu twitter: https://x.com/yule_gan
👉 learn more about neural thickets here: https://thickets.mit.edu
👉 and here with the paper: https://arxiv.org/abs/2603.12228
👉 Slide deck from Yulu: https://www.dropbox.com/scl/fi/ibgfbjmgzdc18t8rvah2g/Neural_thickets_yaccine_interview.pdf?rlkey=uofwkmmytd6idzllqz48woasy&e=1&st=y8t9bkoo&dl=0

Check out Arcee (American open research lab and sponsor of this interview):
♥️ https://www.arcee.ai/
Also trinity large thinking is available on open router:
♥️ https://openrouter.ai/arcee-ai/trinity-large-thinking#providers

Also also for beginners:
📌 learn to code from full-stack to AI with Scrimba https://scrimba.com/?via=yacineMahdid (extra 20% off pro with my link, great resource, I love the team)

Table of Content:
00:00:00 - pre-trained LLM are surrounded by tasks experts
00:05:35 - introduction to the great yulu gan and his beautiful paper
00:08:21 - what are neural thickets?
00:13:25 - relationship between scale and neural thickets?
00:14:56 - why it’s happening for the pretrained base model?
00:18:00 - solution density for small and large model
00:19:30 - do all perturbation help in the same way? and why?
00:25:00 - why does mixed data pretraining create neural thickets?
00:30:20 - would a distill model exhibit a different kind of distribution landscape?
00:32:50 - RandOpt algorithm
00:36:50 - where are the gains coming from with adding random noise?
00:45:45 - what happens to the distribution of the weight when doing the distillation?
00:48:55 - does the useful perturbation of the layers evenly distributed?
00:51:40 - how should we releasing model with a weights distribution perspective?
00:53:56 - does routing work better than distilling you think?
00:56:10 - what do you think will happens at 400B?
00:58:10 - how does this connect to continual learning?
01:02:20 - connection with the baldwin effect from biology?
01:06:30 - discover novelty by random exploration of the weights?
01:09:00 - what’s next for the great yulu gan

enjoy folks hope this stuff is useful🌹

Discussion about this video

User's avatar

Ready for more?