0:00
/
Generate transcript
A transcript unlocks clips, previews, and editing.

How to Generate Synthetic Pretraining Data? with Joël Niklaus from Hugginface

amazing interview with a generous research from hugginface

today we are taking a deeper look at of the secret behind LLM pretraining which is synthetic data pipelines with Joel Niklaus machine learning engineer at Hugging Face!

joel an his team ran 90 controlled experiments and burned over a trillion tokens to figure out what actually makes good pretraining data and in huggingface fashion provided all of their findings/artifacts openly!

lots of things to learn from experiment like this and to be honest there is still so much more to do!

couple of useful links:

👉Joel Website: https://niklaus.ai
👉 Twitter: https://x.com/joelniklaus
👉 Paper : arXiv: https://arxiv.org/abs/2604.13977
👉 Blog: https://huggingface.co/spaces/HuggingFaceFW/finephrase

Also check out Arcee (American open research lab and sponsor of this interview):
♥️ https://www.arcee.ai/
Also trinity large thinking is available on open router:
♥️ https://openrouter.ai/arcee-ai/trinity-large-thinking#providers
♥️ and check out the NAC harness https://www.arcee.ai/blog/nac

on the menu:

- 0:00 : synthetic data in pre-training LLM

- 3:36 : Joel Niklau ML Research Engineer at Hugging Face

- 4:50 : how to approach research problems at the edge?

- 7:26 : the label here is still wrong is should be “benchmarking work overview”

- 9:10 : what are the first few steps to get started building a private benchmark.

- 11:18 : synthetic data pipeline and it’s impact in benchmark making

- 13:19 : Synthetic Data Playbook Presentation Overview

- 19:30 : what is rephrasing in synthetic data?

- 21:05 : why are large scale ablation studies on data pipeline becoming feasible now?

- 24:00 : how is rephrasing done?

- 25:42 : experimental results

- 27:58 : why does synthetic restructuring hurt common-sense knowledge?

- 30:10 : what’s the value prop behind doing rephrasing?

- 32:16 : is math a type of rephrasing?

- 38:02 : what’s up with closeness between generator vs the student?

- 39:40 : synthetic data + mixing real data

- 41:20 : what drive this real-to-synthetic ratio?

- 43:50 : what other structured formats could be explored for synthetic data?

- 45:20 : is the order of datasets used for pretraining also affecting the distribution?

- 46:40 : would adding multilinguality be considered a form of rephrasing for LLM?

- 50:40 : which non-synthetic dataset should you mix in?

- 1:03:20 : do model-specific writing habits survive the rephrasing process? + sawing off the edges?

- 1:08:30 : would their be a gain to have a purpose built 270M model for being a generator?

- 1:14:30 : does generator diversity still help as the student model scales?

- 1:26:30 : ❤️❤️❤️

- 1:27:15 : useful of synthetic data for post-training?

- 1:28:50 : what’s next for this direction and for the field?

btw: do follow joel he is a fantastic scientist and a very good educator!

Discussion about this video

User's avatar

Ready for more?