0:00
/
Generate transcript
A transcript unlocks clips, previews, and editing.

How to Measure Frontier LLM? with Florian Brand from Prime Intellect

all around great presentation by florian of the subject and very pragmatic

Benchmarking LLM in 2026 is now a high stake task that requires deep expertise on LLM architecture, Prompt Engineering, Harness, Sandbox, Scoring, Hardware, Serving, and more. It’s a much different exercise than in the past that have a big impact on the performance of the final model.

Not only that but benchmarking also is starting to have deep cybersecurity concern as frontier models have gained the capabilities of chaining together exploit of various kind (in order to cheat on benchmark).

Underelicitation, Safety concern, cost, all plays a role in figuring out if the new models are the right tool for the job and has downstream consequence across the economy.

Well in this tutorial we are exploring LLM evaluations with a researcher who spends his days taking benchmarks apart: Florian Brand!

he wrote benchmark deep-dives, co-authors the open-models coverage at Nathan Lambert's Interconnects, and now works on evals as a research engineer at Prime Intellect.

we'll go through a gentle intro to why evals matter for beginners, dig into why two scores on the "same" benchmark often aren't comparable, and explore where evals are heading in the age of agents + a live walkthrough of the evaluation module in the Prime Intellect lab!

Pretty packed session, but highly informative!
Enjoy 🌹

important links:
👉 florian twitter: https://x.com/xeophon
👉 prime intellect twitter: https://x.com/PrimeIntellect
👉 learn more about evaluation at prime intellect here:https://www.primeintellect.ai/blog/hosted-evaluations

Check out Arcee (American open research lab and sponsor of this interview):
♥️ https://www.arcee.ai/
Also trinity large thinking is available on open router:
♥️ https://openrouter.ai/arcee-ai/trinity-large-thinking#providers

Also also for beginners:
📌 learn to code from full-stack to AI with Scrimba https://scrimba.com/?via=yacineMahdid (extra 20% off pro with my link, great resource, I love the team)

Table of Content:

0:00:00: are AI Benchmark broken?
0:05:45: Florian Brand background
0:09:00: what motivates florian to work on evaluation?
0:13:33: what is the mirrorcode benchmark about?
0:18:20: cheating in agent benchmark is insaneeeee
0:24:08: LLM benchmarks in era of agents
0:26:30: what’s up with the pelican man
0:28:27: evals are about capabilities
0:31:46: components of running evals
0:35:30: the volume of things to audit is huge!!!
0:40:20: expert answers are wrong hahahaha
0:46:00: api providers aren’t the same
0:48:00: benchmark narrow capabilities (synthetically)
0:50:56: link between eval and environment
0:53:45: small validated benchmark or massive bench?
0:56:11: what is your flow to review a benchmark?
0:58:30: tracking work capabilities with evaluation
1:00:20: slide deck in industry is all vibecoded
1:03:30: harness impact in the evaluation
1:07:39: hardware/sandboxes impact evaluation too!
1:11:00: “is it going to get worse?”
1:12:40: all components influence the final score
1:13:50: training models on different harnesses?
1:17:20: is the model just the weights or it’s all of it?
1:19:30: how to craft benchmark that prevent to cheating and undereliciting models in 2026
1:23:19: ways agents cheat and steal
1:26:00: correct elicitation of capabilities is important
1:36:00: building evaluation on prime intellect
1:45:10: how do you design interactivity benchmarks?
1:48:40: do you think evals are well set to reflect real world performance?
1:52:50: what will the benchmarking landscape will look like in 1 year

Papers & references to check out:
📌
why benchmarking is hella hard: https://epoch.ai/gradient-updates/why-benchmarking-is-hard
📌 b
enchmarking in 2026 situation: https://florianbrand.com/posts/benches-2026

enjoy folks hope this stuff is useful and shout out to my subscribers you guys make doing this at least 300% more fun🌹

Discussion about this video

User's avatar

Ready for more?