Skip to content
All notes

Technology

What are we really evaluating?

On benchmarks, useful systems, and the distance between the two.

2 min readSample note

Placeholder article for the site preview. The text and publication date are illustrative, not published writing by Ian.

A single score is a convenient way to compare complicated systems. It is also a very effective way to hide the details that matter.

When we evaluate a system, we choose not only the tasks but also the definition of success. Those choices embed assumptions about users, environments, and acceptable mistakes.

Look past the average

An average can obscure an uneven experience. A system might perform reliably on familiar requests and fail on an unusual but important case. Looking at errors individually often teaches more than another decimal place in an aggregate metric.

A useful evaluation is a starting point for investigation, not a substitute for judgment. The question is less “Which system wins?” and more “Under what conditions can we rely on it?”

Read next: Leaving a little room for detours