Skip to main content

UX ResearchThink First

Putting the user back into AI Evaluation

  –   The estimated reading time is 10 min.

In our last episode, three researchers described rebuilding their workflows from the ground up. This time we turn to the problem that shows up the moment AI starts producing outputs: evaluation. AI can now generate summaries, themes, documents, even recommendations faster than ever – but how do you know if any of it is actually good?

To find out, we spoke with three researchers at Microsoft AI who have made that question their day job. Christopher Monier and Chuck Kwong are both Principal UX Researchers, and Wendy Wang leads the Copilot UX Research team. The three of them built a response-quality evaluation program together – the kind that compares Copilot’s answers against other AI products and traces the differences back to what users actually value.

Their shared conviction is easy to state but harder to operationalize: machine evaluations can tell you whether an answer meets the criteria you set, but only users can tell you whether it actually worked for them. Putting that user voice back into evaluation is, they argue, the work.

When “good enough” wasn’t obvious

Each of them noticed the evaluation problem from a different starting point.

“When I started at Microsoft AI last year, I didn’t even know what evaluations were,” Chuck says; he came to evaluation as a newcomer. It began with a simple question – “why don’t we see the product the way users do?” – and a simple study – people compared Copilot’s responses with other AI products using the same prompt, and talked through what felt good, bad, useful, or frustrating. The result was a light-bulb moment. Users sometimes strongly preferred one response over another for reasons the team hadn’t been explicitly measuring. “The gap between what teams thought was a good response and what users actually preferred was much larger than expected.”

“A gap between this product and that product, or this model and that model” – that’s what Chris kept sensing without being able to put his finger on it. He had an ambiguous feeling that “something was off” and a suspicion that applying the principles and tactics of UX research could help explain that ambiguity. This itch is what they built UX evaluations to scratch.

Wendy saw it from inside the build. Everyone making AI carries priors about what a good response looks like: its shape, tone, personality, how it should continue the conversation, how closely it should track the prompt. Those priors come from our own experience with AI and our assumptions about how most people want to engage with it. Qualitative conversations with real users kept complicating them. Quality depended on the user’s context and on the goal behind the prompt that never made it into the prompt; content the team would have called off-topic was sometimes the most useful part of the answer; a single phrase, or a callback to an earlier conversation, could change a user’s emotional register. So the work became defining “good” precisely — but defining it from human context and user perspective rather than from what builders already believed.

The dimensions that matter – and who’s equipped to judge them

If evaluation is about measuring quality, the hard part is deciding which dimensions of quality to measure – and which ones a machine simply cannot.

Chris derives his dimensions from those same qualitative conversations. Listening, not just for what people say is important, but reading between the lines for the underlying values that people truly care about. One that surfaced repeatedly: while users want AI to save them time in creating documents, they ultimately want to preserve their point of view and are deeply worried about producing “generic work products.” He typically lands on five to eight dimensions – not generic quality checks – and prioritizes user evaluation for the things the users themselves are uniquely equipped to judge. Accuracy is one example: an LLM can mechanically check whether a claim is factually correct, but only a user can tell you whether, in the context of their company’s strategy and philosophy, that claim is actually relevant. “We can have one AI judge another AI’s output on whether it’s factually accurate,” he says, “but we can’t have one AI judge the other on whether it matters to a person in the real world. Almost by definition, you can’t.”

We can’t have one AI judge another AI’s output on whether it matters to a person in the real world. Almost by definition, you can’t.”

— Chris

For Wendy, the work extends well past the eval itself. Once the data is in, her team turns it into golden data sets: concrete anchors for what a good response and a bad response look like, which then seed much larger machine-run sets.

Just as important is shared vocabulary. The biggest unlock, she says, was building a loss-pattern taxonomy. Describing a response as “generic” or “too shallow” can be accurate without being “a shared language that different teams, from data science to engineers to PMs, could align on.” Once every named loss pattern meant the same thing to everyone, it became “a huge unifying force.” With evals, UXR doesn’t stop at “here are the findings.” It stays in the loop – through prompt iteration, data-set building, and eval design – as the bridge to the user.

The researcher’s unique angle

We asked each of them whether researchers really do bring something to AI evaluation that data scientists and technologists can’t. All three said yes, but for different reasons.

Chuck framed it as a missing voice rather than a turf claim. “I don’t think any one discipline should own or lead evaluations – it’s really a multidisciplinary effort.” But a lot of AI teams build frameworks that lean on LLM-as-a-judge or expert annotators, and “what’s really missing is the user’s voice.” Researchers also bring the measurement expertise to turn subjective concepts like trust, usefulness, or perceived intelligence into signals teams can actually evaluate. That, he says, is exactly what researchers are good at supplying: getting into the room and capturing “what good looks like from the user’s point of view.”

“What’s really missing is the user’s voice.”

— Chuck

Chris agreed “100%” – a data scientist could probably do this work; there’s nothing magic about researchers as people. What’s distinct is the method. Human-centered evals “inherently incorporate all the things that make people human – the imperfections, inconsistencies.” His favorite example: have users write their own prompts. Most people, he notes, are simply not good at prompting – they’re “tired, busy, distracted,” and they leave out context that matters. “If we just have a bunch of AIs generate perfect prompts and run evals on those, we’re missing the reality of how humans interact with AI.” The only way to close “the gap between what the user put in the box and what’s actually in their head” is to systematically ask real people to use the thing and measure whether it lived up to what they expected.

Wendy pointed to a gap in how evals get built. In the push to be rigorous, teams can over-systematize — single-turn, fixed judges, the same controlled set of queries — and in doing so “dilute the type of learnings we can get, because we’re not getting that rich human insight into the whys.” Her fix comes straight from research practice: treat the conversation as the stimulus. “You just think of a conversation as a prototype, as a concept” — dissect it into components and let it drive the learning. That’s what led her team to build multi-turn evals, where people interact with the AI and evaluate the whole exchange rather than a single response.

“You just think of a conversation as a prototype.”

— Wendy

Is the field converging?

Is everyone converging on a shared approach, or building their own? A bit of both, they agreed – with a shared caveat about shelf life. Chuck sees teams still “heads down” building their own frameworks, but sees a measurement science beginning to emerge at the intersection of AI evaluation, UX research, and psychometrics – one focused on systematically understanding how users perceive, experience, and derive value from AI. Chris frames it as best practices and benchmarks that genuinely emerge but never hold still: “just because an evaluation framework was really helpful six months ago doesn’t mean it’s relevant today,” because the frontier keeps moving and users’ expectations move with it. Wendy was more pointed about shelf life: a framework “becomes stale in three months because taste for AI has already evolved.” The convergence worth pursuing, then, isn’t a system. Any system will be outdated by next quarter, but the practice of staying calibrated to where user taste is heading. That applies to her own taxonomy too. Loss-pattern definitions get revised and new patterns get added as expectations move, so the shared language stays current rather than becoming another artifact that describes last quarter’s problems.

Evaluating a model vs. evaluating an experience

A recurring theme was that “evaluating AI” is really several problems under one name.

Chuck drew the sharpest distinction – between evaluating a model and evaluating the whole product experience the model sits inside. The latter is much harder, he says, because of “a lot more confounding variables”: the pixels around the answer, the onboarding journey, the other features, the multi-stage flow in which the model is just one part. Agentic workflows make it harder still – the field has “largely not agreed on how to properly evaluate that entire workflow,” and some stages a user never even sees still shape the experience downstream. Evaluating the model alone is more tractable and more actionable; evaluating the experience “deserves a more thoughtful approach,” with careful study design to isolate what’s actually causing a problem.

Chris highlighted a trend that is fundamentally changing how experiences are evaluated: as AI is introduced into more and more products, its probabilistic nature means “the diversity of user experiences people can have explodes exponentially.” Evaluating the breadth and depth of these experiences is crucial to helping teams understand why metrics like retention and engagement move. And he’s optimistic about the discipline of UX research, pointing out that why people actually use something “is the core of what a UX researcher does,” – and more diverse user experiences simply means more of those questions to answer.

For Wendy, the distinction shows up in what an eval is for. Scoring a response tells you whether this answer was good. The same data, read differently, tells you what the person was actually trying to do, and that’s where the product ideas are. Sometimes they want a quick answer and nothing more. Sometimes they’re working through a hard situation in their life. Sometimes they’re trying to unblock one step in a complicated workflow. Those are different jobs with different definitions of success, and seeing them in the data produces new thinking about what the product should be, not just token-level fixes to the response.

Coming up next

In this episode, across three very different vantage points, the same figure keeps returning: the user, whose answer to the question of what’s genuinely good keeps evolving along with the technology itself.

In the next episode, we continue with another group of researchers and another corner of the practice being remade around AI.

English (United States)
Your Privacy Choices Opt-Out Icon Your Privacy Choices
Consumer Health Privacy Sitemap Contact Microsoft Privacy Manage cookies Terms of use Trademarks Safety & eco Recycling About our ads