Watching five people use the thing you built
Debrief after the sessions, Thursday
Researcher Five done. Four of them landed on the same wall, and it is the wall we argued about in April.
Designer Which wall.
Researcher They open the settings looking for a folder. There is no folder, there are tags, and their mental model wins every time.
PM Four out of five is not a number I can take to anyone. What was task success on the second one.
Researcher Three of five, and time on task doubled for the two who finished. But I would look at the recordings before the number.
Designer The second participant went silent for a whole minute and I still do not know what she was reading.
Researcher That is on me. I nudged her back into thinking aloud twice, then let it go.
PM Who were these people, though. Last round the screener let in three power users and we learned nothing.
Researcher Tightened it. Two had never opened the product before Tuesday.
Designer Can we run the next one unmoderated and get twenty overnight.
Researcher For the click path, yes. For the silence, no — nobody records the pause for you.
PM Then get the quotes on the wall tomorrow, and let us come back to what people were hiring this screen to do in the first place.
An hour of this and the recording is still not the argument. The argument is which of the words in the room means what, and who gets to decide when they collide.

A research round runs from who came to what they were after
The early words are about the room — who was invited, how they were asked, what the moderator did with the silences. The later ones are what survives the room: two numbers, a wall of quotes, and one sentence about what the person actually wanted.
Usability testingTest de usabilidad
Watching people attempt real tasks with a product to find where they get stuck.
Usability testing on Thursday, five participants, same tasks as last round.
The hour itself. Everything else the room argues about is downstream of it.
Screener
A short questionnaire that filters volunteers down to the participants who match the study.
The screener let in too many power users again.
Where a round is won or lost, a week before anyone sits down. Blame for a useless session usually belongs here.
Think-aloud protocolPensar en voz alta
A test technique where the participant narrates their thinking while working, so you hear the reasoning behind the clicks.
She kept going quiet, so we reminded her about the think-aloud protocol.
The moderator’s whole job, said in three words. When it stops, the recording keeps the clicks and loses the reason.
Leading questionPregunta dirigida
An interview question phrased so that it suggests its own answer, which quietly corrupts the finding.
'Wouldn't one button be easier?' is a leading question — ask what they did last time instead.
The way a session quietly turns into agreement with whoever is running it. Everyone knows the rule and everyone breaks it around minute forty.
Task success rateTasa de éxito en la tarea
The share of people who complete a given task without help.
Task success rate went from sixty to eighty-five after we renamed the tab.
The number that travels outside the team, which is exactly why it gets quoted without the sample size.
Time on taskTiempo por tarea
How long it takes to complete a task, used as a usability measure.
Time on task halved, and nobody called support this round.
Sits next to it in every report and says something different: one is whether, this one is at what cost.
Mental modelModelo mental
What a person already believes about how something works, which an interface either matches or fights.
Their mental model is folders, we shipped tags, and that's the whole confusion.
The reason the same screen is obvious to the team and opaque to everyone else. When it collides with the interface, the interface loses.


Unmoderated testingTest no moderado
A study where participants complete tasks alone while a tool records them, with no researcher present.
We ran it unmoderated overnight and had twenty sessions by morning.
What a team reaches for when it needs many more sittings than a week of moderating allows. Cheap on numbers, silent on reasons.
Affinity mappingDiagrama de afinidad
Grouping raw research notes into clusters until themes emerge from the data itself.
We spent the afternoon affinity mapping the interview quotes into six themes.
The afternoon between raw quotes and something a team can act on. Skip it and every session ends in whoever remembers the loudest moment.
Heuristic evaluationEvaluación heurística
A review in which specialists check an interface against known usability principles instead of testing with users.
Before we recruit anyone, let's do a heuristic evaluation and catch the obvious ones.
Done before anyone is invited, to keep real sessions off the obvious problems. It is not a substitute for people, and it gets used as one.
Jobs to be done (JTBD)JTBD
A way of framing a product as the job a person hires it to do, rather than as a feature set or a user type.
She reframed the spec around jobs to be done, and half the feature list stopped making sense.
The question the room comes back to when the findings contradict each other. Confused with a persona often: one is who, this is what for.

There will be another round in six weeks
Same words, a different screen, and probably the same argument about whether three out of five means anything. The one holding the number will be you.
Questions and answers
What about the rest of the product words?
On the deck page, all 150. These eleven are the ones a single research debrief uses.
Why is there Spanish on the card?
Half of these arrive in English wherever the team sits — "screener", "think-aloud" — and get argued about in Spanish. Both forms sit on one card, so the other half of the sentence is not a guess.
Where do the definitions come from?
From the Product & UX deck — the same cards as in the app.

