All Work
  • Measurement Design

Validating an AI Adoption Survey Through Cognitive Interviews

A 20-question survey on AI adoption in UX was nearly ready to launch. Cognitive interviews revealed that one question allowed only positive responses and surfaced five other categories of measurement issues before any data was collected.

Context
DePaul University · 10 weeks · 2025
Role
Lead researcher. Planned and moderated cognitive interviews, led analysis, and translated participant feedback into survey revisions and final recommendations.
Methods
Cognitive interviewing, think-aloud protocol, survey design
Output
6 issue categories corrected before launch
Problem

One question shaped the answer.

The survey was designed to measure AI adoption in UX: frequency of use, trust in outputs, workplace concerns. One question asked only about the most significant benefits of AI in UX. That framing told participants what kind of answer was valid before they responded.

Other questions had different problems. Terms including “transparency,” “consistency,” and “human oversight” appeared without definitions, and participants interpreted them differently.

Two others combined separate ideas into one question, making a clear response impossible. The null response also changed across related items: “Nothing” in one place, “None” in another.

None of these were obvious from inside the team. The survey had been through multiple rounds of internal review and looked finished. Most issues could be corrected with targeted revision. The leading question required a deeper change to the survey logic.

Stakes

Biased questions create invalid findings.

Findings can look rigorous while still reflecting assumptions built into the questions. Once responses are collected, those problems are expensive to correct because revising the instrument can mean starting over.

A participant who held a mixed view of AI in UX had no room to express it. The data would have reflected what the question permitted, and every conclusion built from those responses would carry that error forward.

Research Strategy

Test interpretation before launch.

When researchers read their own questions, assumptions can disappear into familiarity. Cognitive interviews with think-aloud protocol make participant interpretation visible in real time.

Participants
8 participants with UX and HCI backgroundsProduct designers, UX researchers, and HCI students
Sessions
~20 minutes eachIn-person and remote (Zoom)
Experience range
0 to 10+ years
  • Complete the survey aloud

    Participants narrated their interpretation of each question and what influenced their answer. Moderators asked follow-up questions when participants went quiet or answered without elaborating.

  • Probe language, not just answers

    Standard probes included: “What does that term mean to you?”, “Did any part of that question seem unclear?”, “Did the response options let you express what you actually think?”

  • Debrief the full experience

    Post-survey questions surfaced patterns question-level probes could not: which sections felt repetitive, which response scales felt mismatched.

Evidence

Six issue categories surfaced.

Most required targeted revisions. The leading question required a deeper change to the survey’s underlying logic.

  • Leading framing

    Participants recognized that the question assumed AI’s effects were beneficial before they had answered. Its framing and response options excluded critical and mixed experiences, creating a threat to validity rather than a simple wording problem.

    Before “What do you consider the most significant benefits of using AI in your UX work?”

    After “What are the most significant effects of using AI in your role? Select all that apply.”

    Sample options included efficiency, enhanced idea generation, increased reliance on AI, reduction of originality, and bias in outputs.

  • Ambiguous terminology

    Terms including “transparency,” “consistency,” and “human oversight” appeared without definitions. When probed, participants gave meaningfully different answers to the same question.

  • Double-barreled questions

    Two items each combined distinct ideas. Participants answered and then qualified, noting that the question seemed to ask two constructs at once. At least one chose an answer that applied to only one clause and said so unprompted.

  • Redundant questions

    Two pairs of questions covered overlapping themes. Participants flagged both pairs without prompting, asking whether the questions were intentionally different or whether they had missed a distinction.

  • Inconsistent response formatting

    Related items used different labels for the same null response. At least one participant stopped and asked whether the two meant different response states. They did not.

  • Limited response options

    In more than half of sessions, at least one participant wanted an option that did not exist. Several selected the closest available answer and said it did not reflect their actual experience.

Recommendation

Revise the measurement system.

The recommendation translated participant confusion into concrete survey revisions, so each issue had a measurement reason for being changed.

Issue Revision Measurement benefit
Leading framing Reframed the question around AI’s effects Allowed positive, critical, and mixed responses
Ambiguous terminology Added inline definitions Reduced variation in interpretation
Double-barreled questions Split each construct into a separate question Made responses interpretable
Redundant questions Clarified overlapping items Prevented duplicate measures from being treated as separate signals
Inconsistent response formatting Standardized wording and formatting Removed unintended distinctions
Limited response options Added participant-named choices Reduced forced-fit answers
Outcome

Survey Quality

The survey expanded from 20 to 22 questions.

Cognitive interviews led to revisions across the full survey. Two double-barreled items were separated, expanding the instrument from 20 to 22 questions. Ambiguous terms were defined, limited response options were expanded, and repeated formatting was standardized before the survey was distributed.

8
cognitive interviews
6
issue categories
20 → 22
survey questions
Reflection

Test while the draft is still open.

Keep: read like a skeptical participant

The most important finding came from participants who did not share the assumptions built into the survey. Cognitive interviews made that disagreement visible before it became data.

Change: test before the draft feels finished

By the time interviews ran, the survey had been through multiple rounds of internal review and looked finished. That finish may have raised the threshold for pushback: participants were more likely to assume they had missed something than to flag the question as flawed.

Next time, I would test earlier, while the survey still feels provisional. A polished draft can make participants assume confusion is their fault rather than a problem with the question.