NeurIPS 2026, Sydney, Australia, December 11 or 12, 2026

Overview

Can we trust AI evaluation? Modern AI systems are judged through benchmarks, aggregate scores, and public leaderboards, yet trust in these evaluations is often assumed rather than demonstrated. An evaluation can be precise but measure the wrong construct, stable on a familiar benchmark but brittle on newly collected data, or impressive on a leaderboard while poorly aligned with real-world decisions. Repeated benchmark use, small perturbations, underreported variance, leakage, and contamination can further weaken the evidence behind evaluation claims. The TAE (Trust-AI-Eval): Can We Trust AI Evaluation? workshop treats evaluation itself as an object of study: what is measured, which assumptions connect a protocol to a claim, how uncertainty and failure modes are reported, and when the resulting evidence is strong enough to guide deployment. By bringing together work on robustness, causal and measurement validity, auditing, judge reliability, and deployment risk, the workshop aims to clarify when AI evaluation results deserve trust and how evaluation practices can become more reliable, transparent, and decision-relevant.

We invite submissions on topics including, but not limited to:

See the Call for Papers for details.

Submissions will be managed through the OpenReview submission site.

Accepted papers will be presented at the in-person poster session.

Important Dates (Indicative)

Paper submission opens: July 30, 2026
Paper submission deadline: August 29, 2026 (AoE)
Review deadline: September 14, 2026 (AoE)
Author notification: September 22, 2026 (AoE)
Final program posted: September 27, 2026
Workshop: December 11 or 12, 2026


Confirmed Speakers & Panelists

Bin Yu
Bin Yu
University of California, Berkeley
Soheil Feizi
Soheil Feizi
University of Maryland
Liming Zhu
Liming Zhu
CSIRO and University of New South Wales
James Bailey
James Bailey
Monash University

Talk titles and panel details will be announced after the final program is confirmed.


Organizers

Hanxun Huang
Hanxun Huang
University of Melbourne
Barbara Tarantino
Barbara Tarantino
University of Pavia
Paolo Giudici
Paolo Giudici
University of Pavia
Xingjun Ma
Xingjun Ma
Fudan University
Eduard Hovy
Eduard Hovy
University of Melbourne
Sarah M. Erfani
Sarah M. Erfani
University of Melbourne

Sponsors