Researcher at NYU, co-creator of LiveBench, focusing on test set contamination and dynamic evaluation.
Micah Goldblum is distinctive in its focus on creating and promoting a live, dynamic benchmark for LLMs, which sets it apart from more static evaluation methods. The author pushes for the importance of objective judging and avoiding dataset contamination, as seen in the announcement of LiveBench. For example, they highlight the high error rate of LLM judges on hard math and reasoning tasks, which demonstrates their emphasis on rigorous evaluation.
Questions in their segments is part of the full report. Sign in to see it on your own profile.
SIGN IN →Rooms they're not in is part of the full report. Sign in to see it on your own profile.
SIGN IN →A shareable image of this reading — the score, the engines it was measured on, and the date.