← All Thoughts
Brand Strategy

How We Use AI to Measure Impact (and Where We Refuse To)

The AI-in-experiences pitch is loud, and most of it misses the point of measurement. Here are the three places AI genuinely helps us measure impact — and the three places we refuse to let it near.

Justin Ng·18 August 2026·6 min read

The experience-design industry has been pitched a lot of AI in the last eighteen months. Real-time facial-expression engagement scoring. AI moderators for panel discussions. Automated post-program "vibe" summaries generated from photos. Sentiment dashboards that turn attendee smiles into engagement percentages. Attendee-level scoring that ranks who mattered most in the room based on how much they moved.

Most of it misses the point.

Not because AI can't do the things it claims to do — some of it works, technically. But because the things AI is being asked to do in impact measurement often confuse inputs to human judgment with substitutes for human judgment. And when the question is whether an experience actually worked — an event, a wellness retreat, a community program, a brand activation — that distinction is the whole game.

We use AI at Mochi Collective. We use it deliberately, in specific places, across every one of our five practices — most of all in Impact Measurement, where the whole question is whether the work landed. And we tell our clients exactly where and why. Here's the honest breakdown of where it earns its keep — and where we refuse to let it near.

Where we DO use AI

Pattern analysis on post-experience interview transcripts

The most valuable measurement input we have is the follow-up conversation. Seven to fourteen days after any program we've designed — an event, a retreat, a community gathering, a brand activation — we speak with three to five participants for twenty minutes each. Semi-structured, transcribed, honest.

Reading and pattern-coding five to fifteen transcripts by hand takes an afternoon. Feeding them through an LLM to surface recurring language, unexpected phrases, and thematic clusters takes four minutes. The AI doesn't tell us what the experience meant. It gives us a fast, structured starting point — here are the seven things people said unprompted more than once, here are the two phrases that appeared in every transcript, here's a sentiment breakdown by participant.

We still read every transcript ourselves. The AI just makes sure we don't miss a pattern because we were tired on transcript number twelve. It's a very fast intern, not a decision-maker.

Unprompted sentiment monitoring in the wild

The most honest signal of whether a program landed is what people say when no one is asking. Participants who fill out the exit survey say what they think you want to hear. Participants who post about the retreat on LinkedIn three days later, or mention the community gathering in a Slack thread, or bring the brand activation up at a dinner unprompted — that's the signal.

We use light-touch AI monitoring across public channels our clients can legitimately observe (public LinkedIn mentions, X, sometimes internal Slack workspaces where the client has given us access) to catch the unprompted mentions and classify sentiment. Not to grade participants. To catch the phrases that keep appearing organically — so we can compare them against the Monday sentence we predicted at brief stage.

The AI is doing text classification at scale. We're doing the interpretation. The two are not the same job.

Open-text summarisation at scale

A large event, retreat, or membership program usually generates 100+ open-text survey responses. Reading them all takes three hours. Getting a structured summary — recurring themes, outlier responses flagged, quotes worth pulling — from an LLM takes three minutes.

We accept the trade-off: minor loss of nuance in the middle, in exchange for hours of human time freed up to focus on the outliers and the specific quotes that carry the memory. The AI compresses. The human interprets. When the two disagree — which happens — we trust the human.

Where we REFUSE to use AI

The final judgment on whether the work landed

This is the hard line.

The output of impact measurement is not a percentage. It's a judgment: "this program did the job it was designed to do, here's the evidence, and here's what we'd do differently." That judgment is a design responsibility. It requires understanding the client, the moment, the room, the goal that sat behind the brief, and the trade-offs the design made. Machines can produce impressive-looking summaries of the underlying data. They cannot make the call.

If we ever hand a client a report where the top-level verdict was AI-generated, we've abandoned our own craft. So we don't. The judgment is always human, always signed by a specific person on our team, and always defensible when the client pushes back.

Photo and video-based "engagement scoring"

The industry loves this pitch. Cameras in the room analyse facial expressions, body posture, and attention direction, then produce an "engagement score" per audience segment, per session, per speaker, per retreat activity.

We refuse to use it. Two reasons.

It's imprecise. Facial detection can't tell you whether someone is leaning forward because they're engaged or because their back hurts or because they can't hear. It can tell you they look engaged. That's a very different question, and confusing the two produces reports that are confidently wrong.

It's mildly dystopian. Grading participants on visible engagement changes how they show up. The presence of "engagement cameras" is felt by the room whether it's disclosed or not. And the programs we're proudest of tend to have moments where the person who mattered most was sitting quietly and not obviously engaged — just paying attention in a way that mattered later. Systems that score that person low miss the whole point.

Participant-level ranking of "who mattered"

Related but distinct. Some AI tools promise to rank attendees by their contribution — based on speaking time, movement in the room, network centrality, or interaction density.

We won't do this. The person who mattered most at any given program is almost never the one who moved most or spoke loudest. It's usually the one who quietly decided something during a closing circle at a retreat, or the one who said seven quiet words to a founder over coffee at a brand activation that changed a downstream decision, or the community member who didn't post but brought two new members to the next session. That's who the program was designed for, and no scoring system detects them.

Ranking participants creates a false hierarchy that a client can act on. That's worse than not ranking them at all.

The principle underneath

The distinction we hold across all of this is simple: AI as an input to human judgment, never as a substitute for it.

When AI compresses volume that would otherwise cost us time to process — transcripts, sentiment mentions, survey open-text — we use it, because the alternative is that we cost the client more and produce the same output. When AI is asked to make the actual judgment call about whether a program worked, whether a participant mattered, whether a moment landed — that's where our craft lives, and outsourcing it is a betrayal of the client's brief.

We're happy to tell any client, in detail, exactly which parts of their impact report were AI-assisted and which were not. That transparency is part of the deal. It's also part of why the report is trustworthy when we hand it over.

What this means for the brief

If you're commissioning an experience — an event, a retreat, a community program, a brand activation — and someone in the room is pitching you a fully AI-driven "impact measurement platform," ask three questions:

  1. Which specific judgments does the AI make, and which does a named human make? A trustworthy answer names the human.
  2. Where does the AI get its training data, and what's the client's data doing there afterwards? A trustworthy answer knows and can tell you.
  3. What does the tool do when the AI verdict disagrees with human intuition? A trustworthy answer says the human overrides. Everything else is a red flag.

If the answers are fuzzy on any of those three, the "AI-powered measurement" is doing more marketing work than measurement work.

Want us to help you name what "worked" would look like?

The Brief Diagnostic is a free 30-minute conversation. We take your upcoming event, community program, wellness retreat, or brand activation, and we work through what signals would actually tell you whether the experience did its job.

You leave with the Monday sentence, the three signals to track, and the three measurement windows to run them at.

No pitch. No deck. Just your brief and the questions we would ask anyway.

Book a Brief Diagnostic Free · 30 min · we'll ask what you're planning