ImageImage
A one-day, invite-only summit providing a first look at the benchmarks and research that will shape the frontier.

October 8, 2026

San Francisco

Learn what will be on tomorrow’s model
cards from the collaborators behind:

Image
Image
Image
Image
Image
STELLA-Bench
SciHarbor
Paperena
Image
ClinSafe
MedPAIR
FutureSim
PhilosophyBench
AgentAbstain
CollusionBench
HalluWorld
Train-to-Test (T²) Scaling Laws
HumanOversight Bench
JudgementBench
Image
Image
Image
Image
Image
STELLA-Bench
SciHarbor
Paperena
Image
ClinSafe
MedPAIR
FutureSim
PhilosophyBench
AgentAbstain
CollusionBench
HalluWorld
Train-to-Test (T²) Scaling Laws
HumanOversight Bench
JudgementBench

Topics

Image
Environment complexity

Evaluating agents in realistic, dynamic, tool-rich environments.

Image

Autonomy horizon

Measuring how far agents can act independently over long horizons.

Image

Output complexity

Scoring sophisticated, verifiable, multi-artifact deliverables.

Speakers

Image

Sara Hooker

CEO and Founder, Adaption Labs

Image

Alex Ratner

CEO and Founder, Snorkel AI

Image

Russell Yang

Applied Scientist II, Microsoft
Project Lead, JudgementBench

Image

Yiyou Sun

Postdoc, UC Berkeley
Project Lead, AgentLE

Image

Steven Dillmann

Project Lead, Terminal-Bench-Science

Image

Annas Bin Adil

CTO, Atella.ai
Co-Lead, STELLA-Bench

Image

Gabe Orlanski

PhD Student, University of Wisconsin-Madison
Project Lead, SlopCodeBench

Image

Nicholas Roberts

Posdoc Fellow, Princeton University
Project Lead, Train-to-Test (T²) Scaling Laws

Image

Kelly Buchanan

Postdoctoral Researcher, Stanford University
Project Lead, Terminal Bench 2.1

Image

Virginia Smith

Associate Professor, Carnegie Mellon University
Faculty Lead, CollusionBench

Image

Vincent Sunn Chen

Founding Team, Snorkel AI

Shaping the frontier of AI

Livestream

Stay updated on
event announcements