0:00
/

Inside a Frontier Open Model: Nemotron 3 Ultra Explained ft. NVIDIA's Chris Alexiuk

today we are taking a deeper look at how you actually build a frontier open AI model with Chris Alexiuk, senior product research engineer at NVIDIA

open models are infinitely interesting to me. being able to peek under the hood and look at the interconnected pieces that each hums and work in a specific way to feed the whole is a blessing.

I had a fantastic time earlier this year chatting with chris about the nemotron family of models and why they are engineered the way they are.

throughout the discussion it became clear all decisions surround nemotron 3 ultra was about making it able to be fast fast and handle long context efficiently.

huge thanks to the nemotron team for putting their stuff out there like that! ❤️

Table of Content

  • 0:00 : nemotron 3 ultra ethos

  • 3:42 : Chris the Senior Product Research Engineer (phew)

  • 6:10 : Nemotron Lab & the Coalition

  • 9:10 : the open model ecosystem party 🤗

  • 12:50 : “a faster model is a smarter model”

  • 20:23 : intro to nemotron family

  • 38:18 : infrastructure issues on the research side for these type of models?

  • 42:30 : main design choices that sets nemotron 3 ultra apart?

  • 48:10 : long context inside the models or inside the harness?

  • 57:35 : Latent MoE tradeoff or is it a free 🥪

  • 1:01:14 : very aggressive grouped query attention why?

  • 1:08:00 : what in the structure of nvidia enabled this whole thing?

  • 1:11:36 : advice for undergrad that want to work on this stuff?

  • 1:15:45 : how does the shared weight MTP work?

  • 1:19:20 : main idea behind MOPD?

  • 1:24:55 : what’s next for nemotron scaling?

  • 1:26:52 : any weird reward hacking stories in nemotron?

  • 1:30:30 : what’s next for nemotron ++++?

Discussion about this video

User's avatar

Ready for more?