The Hidden Influence of AI Benchmarks

July 24, 2026

Author: Derek Chezzi    Editor: George Stein

How outdated evaluation metrics in AI benchmarks can steer innovation away from what people actually value.


Imagine two AI image-generation models side by side.

One produces images that look more realistic to human viewers. The objects make sense. The scene holds together. The image feels coherent in the way people intuitively judge visual realism.

The other model receives the better benchmark score.

That tension sits at the centre of a larger question facing AI development: are the tools we use to measure progress still measuring the kind of progress that matters?

In a now widely cited study from our lab, we set out to test whether the benchmarks that guide the development of today’s image-generation models still capture what people actually value as improvements. We wanted to know if the most widely used evaluation methods were keeping pace with the capabilities of modern systems, or if they were rewarding optimizations that no longer mattered most to users.

This is not just a technical issue.

In AI, benchmarks are more than scoreboards; they quietly shape the direction of progress itself. They influence what researchers optimize for, what companies promote, and what the field recognizes as improvement. When those benchmarks fall out of step with human judgment, the risk isn’t simply that some models are ranked incorrectly. The deeper risk is that the field begins moving toward the number, rather than toward better performance for people.

That realization guided our evaluation study. We wanted to understand whether the metrics that helped early models improve still reflect what today’s systems can do, and whether a more modern approach could realign those measurements with human perception.

Generative AI has changed. Evaluation has to change with it.

Earlier AI-generated images were often easier to judge. If a model produced a blurry object, a distorted face, or an obviously artificial scene, both people and metrics could generally agree that the output was poor.

But image-generation models have improved quickly. As outputs became more realistic, the differences between models became more subtle, and the evaluation problem changed.

At that stage, the question is no longer simply whether an image is sharp, colourful, or detailed. The question becomes whether the image makes sense as a whole.

That is where older image-quality benchmarks can begin to show their limits. 

A benchmark that worked well when the field was separating obviously poor images from better ones may be less useful when the task is distinguishing between high-performing models. Once generated images start to look convincing to people, evaluation tools need to be sensitive to the same kinds of qualities people notice.

This is especially important because benchmarks do not sit outside the development process; they become part of it. Researchers use them to compare approaches. Companies use them to support claims of progress. The broader AI community uses them to decide which models are improving and which directions are worth pursuing.


So, when an evaluation tool is misaligned, it can do more than produce an imperfect ranking. It can create imperfect incentives. By turning evaluation metrics into optimization targets, the field has inadvertently triggered Goodhart’s Law:

When a measure becomes a target, it ceases to be a good measure.

Goodhart’s Law

What we wanted to understand

Our research began with a simple, practical question that had surfaced across the AI community. We wanted to answer it with data rather than intuition.

If newer image-generation models are producing images that look more realistic and coherent to people, why do some of the most widely used benchmarks still rank them as “worse”?

Many researchers had already suspected that certain evaluation metrics might be falling out of step with human judgment. But suspicion alone doesn’t change practice. When a benchmark is widely used, it continues to shape the field until clear evidence shows where it falls short and what might work better.

We set out to test that alignment directly. We compared how commonly used image-evaluation metrics ranked generative models against how human evaluators judged the realism of the same images.

At a high level, our study focused on three questions:

  1. Alignment: Do existing benchmarks reflect human judgments of image realism?
  1. Mismatch: If not, what explains the disconnect?
  1. Improvement: Can a more modern evaluation approach capture the qualities that people actually see as realism?

The answers matter because evaluation is not just about producing a leaderboard. It determines what kind of model behaviour the field rewards, and ultimately, what kind of progress it makes.

What we found

The clearest finding was that human judgment and benchmark rankings can diverge.

In our experiments, commonly used image-generation metrics did not always strongly align with human evaluations of realism. That gap matters because diffusion models—the systems behind many of today’s most realistic image tools—often look better to people even when older metrics don’t capture it.

Figure 1: Models are ordered by FID, while the y-axis shows human error rate (how often people mistook generated images for real ones). A strong alignment between FID and human judgment would produce a clear trend. The absence of that trend shows where benchmark rankings and human perception diverged.

That raised an important question: what was the metric actually measuring?

Part of the issue came down to how the benchmark represented images internally. Older approaches emphasized local textures more than global coherence, the sense that the scene as a whole makes visual sense.

A person looking at an image may quickly recognize whether the scene makes sense. Does the animal have the right structure? Do the objects belong together? Does the image hold together as a believable whole?

A metric may be sensitive to different things. It may detect patterns that are meaningful mathematically, but less connected to the way people judge visual realism.

Our experiments showed that more modern evaluation methods aligned more closely with human judgment, reinforcing that how a metric is designed shapes what progress looks like.

What this means for image evaluation

The immediate takeaway is that evaluation metrics are not neutral.

Every benchmark makes choices. What counts as quality? Which differences matter? What should be rewarded? What gets compressed into a single number?

Those choices become more important as models become more capable.

This does not mean every older benchmark should be discarded. It also does not mean a newer metric should be treated as the final word. But it does mean the field needs to keep questioning whether its evaluation standards are keeping pace with the systems they are meant to assess.

That is especially true when one score is used to stand in for overall model quality.

A single number is useful because it simplifies comparison. It helps researchers and companies answer a question everyone wants answered: which model is better?

But “better” is rarely one-dimensional. A model might produce more realistic images but less variety. Another might generate a broader range of images, but with more visual flaws. Yet another might appear realistic because it is reproducing training examples too closely.

Those distinctions matter. If they are collapsed too casually into one score, the score may become easier to compare but harder to interpret.

Benchmarks shape what the field optimizes for

This is where the issue widens beyond the scope of image generation. Across the board, benchmarks do not simply measure progress. They help define it.

Benchmarks are necessary because they give the field shared reference points. But as they become standards, they become targets—researchers optimize toward them, and companies cite them. The problem emerges when those targets no longer reflect the kind of progress that matters most. 

If a metric rewards improvements that do not translate into better user experience, teams may spend time and money improving the number without improving the model in ways people can see or use. In image generation, that may mean optimizing for a score while downstream users do not experience a meaningful improvement in output quality.

The broader lesson is that measurement choices become development choices. If we reward the wrong thing, we should not be surprised when the field gets better at producing this “mistake.”

Why changing benchmarks is difficult

Even when better evaluation methods exist, changing a benchmark is rarely straightforward. The inertia usually comes down to three things:

  • History: Older metrics make it easy to compare new results against years of past research.
  • Practicality: If a new benchmark is too complex, researchers won’t adopt it for daily experimentation.
  • Incentives: If an older metric makes a team’s results look stronger, they have little motivation to switch to one that reveals weaknesses.

Furthermore, any metric must balance inherent trade-offs. Image evaluation usually considers fidelity (realism), diversity (variety), and memorization (reproducing training data rather than generating new data). Because improving one can make another worse, a benchmark that hides these trade-offs behind a single number can obscure the actual progress being made.

That is why benchmark reform is never purely a technical issue. It is also a matter of standards, incentives, and communication. The field still craves a single, simple ranking —one score to rule them all—yet simplicity often comes at the cost of nuance. The goal is not to abandon benchmarks altogether, but to be more thoughtful about what they leave out and how those omissions shape the path of progress.

The cost of not changing

The risk of outdated benchmarks isn’t dramatic failure, but misdirected progress. When the field optimizes for scores that no longer reflect human judgment, teams may spend time and resources improving numbers that don’t translate into better results for users. A model can climb a leaderboard without offering a clearer, more coherent, or more inventive experience.

For image generation, that misalignment has practical consequences. Effort that could strengthen a model’s understanding of structure, composition, or realism might instead go into fine-tuning metrics that reward superficial improvements. Over time, progress starts to look impressive on paper but less meaningful in practice.

For AI more broadly, the lesson is simple but far-reaching: as systems become more capable, evaluation methods must evolve with them. Otherwise, the field risks mistaking movement for progress.

Our research doesn’t claim that all benchmarks are flawed. It simply shows that widely adopted tools can drift away from human judgment, and that newer approaches are needed to bring measurement back into line. The next generation of AI progress will depend not only on building stronger models but on asking sharper questions about how they are judged. Better measurement does more than record improvement; it ensures the field moves in a direction that people actually recognize as progress. 

How We Tested Image-Generation Metrics Against Human Judgment

Our research examined whether widely used image-generation evaluation metrics align with human judgments of image realism.

At a high level, we compared outputs from a wide range of generative image models, looked at how existing metrics ranked those outputs, and then compared those rankings with human evaluations.

For years, researchers have relied on metrics such as Fréchet Inception Distance (FID), Inception Score, and related measures to compare image-generation models. Many of these metrics follow a similar two-step structure. They use an encoder to convert images into a lower-dimensional representation, and then calculate a distance or comparison between real and generated samples within that space.

That encoder matters. In the case of FID, the standard representation comes from Inception-V3, a network trained for classification on ImageNet1k. The underlying assumption is that this representation captures features that are meaningful for evaluating image quality across all domains.

However, questions had already been raised about whether Inception-based representations are too limited for modern image-generation evaluation. Researchers have suggested that they may put too much emphasis on texture and object-class cues rather than broader perceptual qualities.

We set out to test that assumption directly. Alongside Inception-V3, we evaluated several modernized alternative representation spaces, including:

  • Self-supervised vision transformers such as DINOv2 and MAE 
  • CLIP-based and contrastive models such as CLIP, SwAV, SimCLRv2, and data2vec 
  • Representations designed for perceptual similarity, such as DreamSim

Our goal was to see whether a different encoder could produce evaluation scores that better reflect human perception.

We tested 41 generative models across four image datasets: CIFAR-10, ImageNet-1k, FFHQ, and LSUN-Bedroom. These models represented six main model families: diffusion models, Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), normalizing flows, transformer-based models, and consistency models. From each model, we generated 100,000 images, for a total of 4.1 million generated images, and evaluated them across 17 metrics.

We also conducted a large human evaluation study. More than 1,000 paid participants were shown either a generated image or a real image from the training dataset and asked whether it was real or fake. Across the study, we collected 207,000 individual responses. To our knowledge, this was one of the largest human-subject experiments on generative model evaluation conducted to date.

Those responses gave us a human baseline for image realism (fidelity). We used human error rate as the key measure. If people had a harder time distinguishing a generated image from a real one, we considered the model to be producing more realistic images.

The core finding was that common metrics did not always align with human judgment. In particular, diffusion models were often judged by people as producing more realistic images, even when older Inception-based benchmark scores ranked them less favourably.

One key issue was the representation space used by the metric. Older Inception-based evaluation could focus more on localized textures, patterns, and class-related features than on overall image coherence. Humans, by contrast, often judge whether an image makes sense as a whole.

Figure 2: Different encoders “see” images differently. Inception-based evaluation tends to focus narrowly on prominent objects or localized visual features, while self-supervised alternatives such as DINOv2 capture a more holistic view of the image. This helps explain why the representation space behind a metric can affect whether its scores align with human judgment.

When we tested alternative representation spaces, we found that FD calculated with DINOv2-ViT-L/14 aligned better with human evaluations than the standard Inception-based approach. This does not make it a final solution to image evaluation, but it showed that changing the underlying representation can improve the relationship between benchmark scores and human judgment.

We also tested alternative explanations for the mismatch. For example, we examined whether models that looked more realistic to people were achieving that by sacrificing diversity and producing a narrower range of images. Our analysis did not support that explanation. We also examined rarity and memorization, including whether models were reproducing training examples too closely rather than generating genuinely novel images.

Together, these findings point to a broader evaluation challenge: image quality is not fully captured by a single number. Fidelity, diversity, rarity, and memorization each answer different questions about model performance. A useful benchmark may simplify comparison, but it should not hide the dimensions that matter.

Why One Quality Score Is Not Enough

A single image-quality score can be useful. It gives researchers and companies a simple way to compare models. It makes leaderboards possible. It creates a shared reference point.

But image-generation quality is not one thing.

A model can perform well in one dimension and poorly in another. That is why relying too heavily on a single score can hide important distinctions. A model may generate images that look realistic but produce too little variety. Another may generate a broader range of images, but with more visible flaws. Another may appear highly realistic because it is reproducing training examples too closely.

Those are different problems, and they require different forms of evaluation.

Figure 3: An illustration of learned distributions and samples (orange, crosses) having different properties with respect to the true distribution and training set (blue, squares). Italicized text indicates metrics that purport to detect these properties.

Fidelity asks: Do the generated images look convincing to humans?

A realistic image is not just sharp, colourful, or detailed. It should make sense as a whole. The object should have the right structure. The scene should hold together. The image should not contain visual errors that make a person immediately recognize it as artificial.

This is the dimension most people intuitively think of when they ask whether an image model is “good.”

Diversity asks: Can the model generate a wide range of outputs?

A model that produces one excellent cat image over and over may look impressive in a small sample, but it is not doing what we want a generative model to do. A stronger model should be able to produce many different realistic images, not just repeat a narrow set of successful patterns.

Diversity matters because image generation is not only about producing one convincing output. It is about capturing the range and variation we expect from the real world.

Rarity asks: Are certain images unusual within the real image distribution, even if they are still real?

A rare image is not necessarily unrealistic. It may simply be less common or less typical than other images in the dataset. This matters because people could, in theory, mistake unusual real images for fake ones. In our research, we tested that possibility and found that human evaluators were not simply confusing rare images with unrealistic images.

Memorization asks: Is the model generating new images, or is it reproducing images it has already seen?

This is more of a check than a quality measure. A copied image can look realistic, but that does not mean the model is generating well. It may simply be reproducing data too close to its training data.

That distinction matters because realism alone can be misleading. A memorized image may score well on appearance, but it raises a different evaluation concern.

Taken together, these dimensions show why “best model” is often too simple a question. Better evaluation starts by asking: best at what, for whom, and according to which measure?

Citation

@inproceedings{
stein2023exposing,
title={Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models},
author={George Stein and Jesse C. Cresswell and Rasa Hosseinzadeh and Yi Sui and Brendan Leigh Ross and Valentin Villecroze and Zhaoyan Liu and Anthony L. Caterini and Eric Taylor and Gabriel Loaiza-Ganem},
booktitle={Thirty-seventh Conference on Neural Information Processing Systems},
year={2023},
url={https://openreview.net/forum?id=08zf7kTOoh}
}