One question shaped the answer.
The survey was designed to measure AI adoption in UX: frequency of use, trust in outputs, workplace concerns. One question asked only about the most significant benefits of AI in UX. That framing told participants what kind of answer was valid before they responded.
Other questions had different problems. Terms including “transparency,” “consistency,” and “human oversight” appeared without definitions, and participants interpreted them differently.
Two others combined separate ideas into one question, making a clear response impossible. The null response also changed across related items: “Nothing” in one place, “None” in another.
None of these were obvious from inside the team. The survey had been through multiple rounds of internal review and looked finished. Most issues could be corrected with targeted revision. The leading question required a deeper change to the survey logic.
Biased questions create invalid findings.
Findings can look rigorous while still reflecting assumptions built into the questions. Once responses are collected, those problems are expensive to correct because revising the instrument can mean starting over.
A participant who held a mixed view of AI in UX had no room to express it. The data would have reflected what the question permitted, and every conclusion built from those responses would carry that error forward.
Test interpretation before launch.
When researchers read their own questions, assumptions can disappear into familiarity. Cognitive interviews with think-aloud protocol make participant interpretation visible in real time.
- Participants
- 8 participants with UX and HCI backgroundsProduct designers, UX researchers, and HCI students
- Sessions
- ~20 minutes eachIn-person and remote (Zoom)
- Experience range
- 0 to 10+ years
-
Complete the survey aloud
Participants narrated their interpretation of each question and what influenced their answer. Moderators asked follow-up questions when participants went quiet or answered without elaborating.
-
Probe language, not just answers
Standard probes included: “What does that term mean to you?”, “Did any part of that question seem unclear?”, “Did the response options let you express what you actually think?”
-
Debrief the full experience
Post-survey questions surfaced patterns question-level probes could not: which sections felt repetitive, which response scales felt mismatched.
Six issue categories surfaced.
Most required targeted revisions. The leading question required a deeper change to the survey’s underlying logic.
-
Leading framing
Participants recognized that the question assumed AI’s effects were beneficial before they had answered. Its framing and response options excluded critical and mixed experiences, creating a threat to validity rather than a simple wording problem.
Before “What do you consider the most significant benefits of using AI in your UX work?”
After “What are the most significant effects of using AI in your role? Select all that apply.”
Sample options included efficiency, enhanced idea generation, increased reliance on AI, reduction of originality, and bias in outputs.
-
Ambiguous terminology
Terms including “transparency,” “consistency,” and “human oversight” appeared without definitions. When probed, participants gave meaningfully different answers to the same question.
-
Double-barreled questions
Two items each combined distinct ideas. Participants answered and then qualified, noting that the question seemed to ask two constructs at once. At least one chose an answer that applied to only one clause and said so unprompted.
-
Redundant questions
Two pairs of questions covered overlapping themes. Participants flagged both pairs without prompting, asking whether the questions were intentionally different or whether they had missed a distinction.
-
Inconsistent response formatting
Related items used different labels for the same null response. At least one participant stopped and asked whether the two meant different response states. They did not.
-
Limited response options
In more than half of sessions, at least one participant wanted an option that did not exist. Several selected the closest available answer and said it did not reflect their actual experience.
Revise the measurement system.
The recommendation translated participant confusion into concrete survey revisions, so each issue had a measurement reason for being changed.
| Issue | Revision | Measurement benefit |
|---|---|---|
| Leading framing | Reframed the question around AI’s effects | Allowed positive, critical, and mixed responses |
| Ambiguous terminology | Added inline definitions | Reduced variation in interpretation |
| Double-barreled questions | Split each construct into a separate question | Made responses interpretable |
| Redundant questions | Clarified overlapping items | Prevented duplicate measures from being treated as separate signals |
| Inconsistent response formatting | Standardized wording and formatting | Removed unintended distinctions |
| Limited response options | Added participant-named choices | Reduced forced-fit answers |
Survey Quality
The survey expanded from 20 to 22 questions.
Cognitive interviews led to revisions across the full survey. Two double-barreled items were separated, expanding the instrument from 20 to 22 questions. Ambiguous terms were defined, limited response options were expanded, and repeated formatting was standardized before the survey was distributed.
- 8
- cognitive interviews
- 6
- issue categories
- 20 → 22
- survey questions
Test while the draft is still open.
Keep: read like a skeptical participant
The most important finding came from participants who did not share the assumptions built into the survey. Cognitive interviews made that disagreement visible before it became data.
Change: test before the draft feels finished
By the time interviews ran, the survey had been through multiple rounds of internal review and looked finished. That finish may have raised the threshold for pushback: participants were more likely to assume they had missed something than to flag the question as flawed.
Next time, I would test earlier, while the survey still feels provisional. A polished draft can make participants assume confusion is their fault rather than a problem with the question.