October 8, 2026
San Francisco
Topics
Evaluating agents in realistic, dynamic, tool-rich environments.
Autonomy horizon
Measuring how far agents can act independently over long horizons.
Output complexity
Scoring sophisticated, verifiable, multi-artifact deliverables.
Speakers
Sara Hooker
CEO and Founder, Adaption Labs
Alex Ratner
CEO and Founder, Snorkel AI
Russell Yang
Applied Scientist II, Microsoft
Project Lead, JudgementBench
Yiyou Sun
Postdoc, UC Berkeley
Project Lead, AgentLE

Steven Dillmann
Project Lead, Terminal-Bench-Science
Annas Bin Adil
CTO, Atella.ai
Co-Lead, STELLA-Bench
Gabe Orlanski
PhD Student, University of Wisconsin-Madison
Project Lead, SlopCodeBench
Nicholas Roberts
Posdoc Fellow, Princeton University
Project Lead, Train-to-Test (T²) Scaling Laws
Kelly Buchanan
Postdoctoral Researcher, Stanford University
Project Lead, Terminal Bench 2.1
Virginia Smith
Associate Professor, Carnegie Mellon University
Faculty Lead, CollusionBench
Vincent Sunn Chen
Founding Team, Snorkel AI
