Shrshnk is a researcher working on AI evaluations and benchmarks. His work focuses on the development of reliable and robust evaluation methods for AI systems.
Shrshnk pushes a distinctive angle by emphasizing the importance of looking at one's own data and not relying on generic off-the-shelf metrics. Shrshnk does this by providing concrete examples, such as the 9 mistakes people make when building evals for their AI apps, and highlighting the limitations of automating evals with AI. Shrshnk's focus on the nuances of AI evaluation and development sets them apart from generic peers, as seen in their post '9 mistakes people make when building evals for their AI apps'
Questions in their segments is part of the full report. Sign in to see it on your own profile.
SIGN IN →Rooms they're not in is part of the full report. Sign in to see it on your own profile.
SIGN IN →A shareable image of this reading — the score, the engines it was measured on, and the date.
CiteGist uses essential cookies to keep you signed in. With your consent we'd also use analytics cookies to understand which pages help most. Learn more.