The keynote ends, the applause fades, and within the hour a survey lands in every attendee's inbox asking them to rate the session out of ten. The scores come back strong, someone pastes the average into the wrap-up deck, and the event is declared a success. This is how most organisations evaluate their single largest speaker investment — and it measures almost nothing.

Here is our position, stated plainly. Satisfaction scores measure whether people enjoyed the hour, not whether anything changed, and the two are far less related than organisers assume. A delighted audience may have learned nothing. A room that laughed in all the right places can file out exactly as it filed in, carrying warm feelings and no new capability whatsoever. The happy sheet tells you whether people had a pleasant hour. It cannot tell you whether anything happened.

The ratings compound the problem by being uniformly flattering. Professional speakers are good at their job — that is why they can charge for it — and audiences are polite, so session scores cluster so high across the board that they cannot tell a good booking from a great one. Nine point something is the standard result at every fee level, which means the score has stopped carrying information. The honest signal, in the market we work in, is blunter and more commercial: whether the client books the speaker again. Repeat business is the only review that costs the reviewer something.

The ladder almost nobody climbs

When our consultants talk measurement with clients, we use a ladder. On the bottom rung sits reaction — did people enjoy it. Above that, learning — can they repeat the ideas. Then application — has anyone done anything differently back at work. Then business impact — did a measure the organisation cares about move. And at the top, return — the value of that movement set against the full cost of the event. Beneath the whole ladder sit the inputs: attendance, cost per head, the things that are easy to count precisely because they say nothing about value.

Almost every event is measured at the bottom of that ladder, and almost none at the top. Organisations count who came and how they felt, then stop — and a favourable reaction does not ensure learning, which is the satisfaction problem restated from the other direction. The rungs that would actually justify the fee sit above the ones anybody bothers to stand on.

Before you commission a control-group study for your next conference, though, hear the second half of the argument. Only a minority of events warrant a full evaluation, and measurement should match the stakes. A quarterly staff briefing does not need an impact study. A sales kickoff built around one strategic message, carrying a serious fee and a serious ambition, probably does. The discipline is choosing your rung deliberately — not defaulting to the bottom one because the survey tool was already configured.

A middle path that costs almost nothing

For most keynotes, the sensible answer sits between the happy sheet and the full study. It starts before the event, with a written decision about what success looks like. Then it pairs an immediate signal with two delayed ones.

The immediate signal can stay simple — a session net promoter score or equivalent, which is really a measure of advocacy: whether people would recommend the experience. Useful, but it captures resonance, not learning. So pair it with a retention check at around two weeks — can attendees actually repeat the speaker's frameworks, unprompted? — and an application check at 30 to 60 days: has anyone done anything differently? Two short pulse questions at each interval will do. This is the cadence we recommend to any client with a serious fee on the table, and it costs two emails.

Ask better questions while you are at it. The most useful reframe we know is to ask the audience how did this help? rather than how did I do? — the first question surfaces evidence, the second solicits applause. And watch for the most honest metric available: who spontaneously repeats the key messages afterwards. When the speaker's language starts appearing in meetings six weeks later, without prompting, you have evidence no survey can fake.

The gap the industry keeps admitting

None of this is controversial, which makes the state of practice all the stranger. Nearly every events team we meet says it wants better measurement, and nearly every one still tracks attendance and satisfaction. The industry says it wants the top of the ladder and keeps standing on the bottom rung — not from laziness, mostly, but because the bottom rung is the only one the standard tools were ever built to reach. The survey platform fires automatically. The two-week retention check requires someone to decide, in advance, that it matters.

You do not need a research department to do better. You need one page written before the event naming the level you intend to measure at, and two calendar reminders — a fortnight out and two months out. The happy sheet will keep coming back at nine out of ten regardless. That was never the question.