Stephan Rabanser is a researcher working with Kapoor and Narayanan on the critique of benchmark accuracy. His work is helping turn the critique into a more formal science of AI reliability—consistency, robustness, predictability, and safety rather than simply 'percent correct'.
Questions in their segments is part of the full report. Sign in to see it on your own profile.
SIGN IN →Rooms they're not in is part of the full report. Sign in to see it on your own profile.
SIGN IN →A shareable image of this reading — the score, the engines it was measured on, and the date.