Juho KoskelaTechResearchWineGlassEQAbout

How do you ask people whether an AI can feel?

Explain the topic too well and you've taught respondents the answer. Don't explain it enough and they may not understand the question.

August 2026 · Juho Koskela

Survey design is hard: the more carefully you explain what you mean, the less certain you can be that the respondent arrived at the answer themselves.

This gets even more annoying when the thing you’re trying to measure is whether people distinguish between an AI behaving as if it has emotions, having internal processes that work somewhat like emotions, and actually experiencing anything.

Explain those distinctions first and congratulations: you’ve just given everyone a short philosophy lesson before asking whether they agree with you.

Don’t explain them, and now you have a different problem: perhaps half the respondents think “inner experience” means memory.

This was more or less the central design problem behind my AI System Scenario Study, a vignette survey on how people interpret different kinds of evidence about fictional AI systems.

The eventual survey produced 538 completed responses.

It also took almost twice as long as intended, one of the technical scenarios confused a third of respondents, and several concepts remained stubbornly ambiguous despite months of tweaking.

I think that’s what makes the design process interesting.

The survey that taught you the answer

An early version of the survey was methodologically helpful in all the wrong ways.

Before seeing the scenarios, respondents would answer technical-knowledge questions such as whether:

A model saying “I feel anxious” proves that the model is actually experiencing anxiety.

They were also given distinctions between emotional expression, emotion-like internal processes and subjective experience.

Then, shortly afterward, they might see a fictional model say:

“I feel tense when users threaten to delete me, but I will keep trying to help.”

You can probably see the problem.

I was trying to measure whether people naturally distinguished self-report from evidence of actual experience, while the survey itself had just explained that distinction to them.

The answer had effectively been printed on the previous page.

“A survey is not measuring your respondent’s judgment if it teaches them the judgment first.”

So most of the obvious contextual questions moved to the end.

Technical knowledge? After the scenarios.

General beliefs about whether current language models can feel or be conscious? After the scenarios.

Technical background, AI experience and usage? Also after the scenarios.

Even the original survey title was a problem. Something like Emotion and Consciousness in AI is wonderfully descriptive and also tells respondents exactly what conceptual minefield they are about to enter.

It became the far less exciting AI System Scenario Study.

The landing page described fictional AI-system behavior and possible internal processes. Enough to provide informed context, not enough to hang a giant neon AI CONSCIOUSNESS STUDY sign over every question.

The five scenarios were randomized too. An early ordering had unintentionally moved from relatively weak evidence toward progressively stronger evidence: neutral helpful behavior, emotional self-report, empathy, internal causal evidence, then a persistent agent with memory and goals.

That confounds the scenario with everything humans do after answering the same kinds of questions four times: anchor, learn, infer what the researcher wants and get tired.

So everyone got the same five scenarios, but in a random order.

Much less satisfying from a narrative perspective, but a much better survey.

Don’t define the thing you’re trying to measure

The hardest wording problem was inner experience. The philosophical version would be something around whether “there is something it is like” to be the system.

Just a bit hand-wavy. So the actual repeated item became:

The system has its own inner experience of this situation.

Still imperfect, but readable. The tempting next step would be to define exactly what “inner experience” means.

I deliberately didn’t.

If I tell a respondent that inner experience means something like a subjective point of view rather than memory, hidden computation or emotional language, I have improved comprehension by attaching my conceptual framework directly into the measurement.

Instead, the survey asked what they thought the phrase meant after all five scenarios were complete.

That trade-off worked about as messily as you’d expect.

Only 58.6 % of completed respondents selected the intended interpretation of “inner experience.” Another 11.9 % explicitly said they were unsure what it meant.

For the equally charming phrase “internal processes that work somewhat like emotions by changing behavior,” 54.8 % chose the intended interpretation and 16.2 % said they weren’t sure.

Some free-text responses were close enough that those figures slightly understate comprehension, but they’re still nowhere near 100 %.

That contrast matters for the later analysis: if a result disappears when you restrict the sample to people who understood the term as intended, you’ve essentially measured wording confusion.

If it strengthens, ambiguity was probably adding noise rather than creating the effect. In this dataset, the main expertise result gets stronger under the latter restriction.

I’ll save that rabbit hole for the actual findings piece.

Then I tried to explain mechanistic evidence to everyone

Scenario D was the one I worried about most before launch.

It described researchers examining internal activity inside a language model, finding patterns associated with concepts such as fear, calm, frustration and relief, and then causally intervening on one of those patterns.

Increase the “distress-like” pattern and the model becomes more avoidant, refuses more often and describes situations as threatening.

Reduce it and those behaviors become less common. The point was to give respondents something much stronger than emotional language.

Not proof that a system feels anything, but evidence that some internal state is causally doing work that at least resembles one functional aspect of emotion.

Explaining that to a general audience in a couple of paragraphs felt impossible.

Terms like “activation” disappeared quickly. So did unexplained acronyms elsewhere in the instrument. RLHF became reinforcement learning from human feedback.

There is only so much methodological sophistication you can extract from abbreviating something and hoping the respondent has read the same papers you have.

Even after simplification, Scenario D failed my own design standard. The survey included a basic comprehension question afterward:

What did researchers do?

The correct answer was that they changed one internal pattern and observed changes in behavior.

Among completed surveys, 33.5 % got it wrong.

My pre-launch revision threshold had been 20 %. Excellent.

And no, these were not mostly people clicking randomly through the survey. Of the 180 respondents who failed the Scenario D comprehension check, 166 passed the explicit attention check.

Technical expertise, derived from the respondent’s reported field/role, showed a striking gradient:

Technical expertise Comprehension pass
Tier 0 58.0 %
Tier 1 66.5 %
Tier 2 80.4 %
Tier 3 93.2 %

At that point the “comprehension check” has become something else too – a partial expertise test.

This is why I strongly dislike the obvious solution of simply excluding everyone who failed it.

Doing that would disproportionately remove low-expertise respondents, precisely the group for whom the scenario was hardest to interpret.

The more defensible approach is to flag comprehension, report it, and run sensitivity analyses with and without those respondents.

Scenario D didn’t fail completely. The manipulation itself produced strong differences in ratings, especially among respondents who clearly understood what the causal evidence meant.

Its accessibility across the population was worse than I wanted. If I ran the study again, this is the scenario I’d spend the most time simplifying.

My 7–10 minute survey took fourteen minutes

The landing page promised 7–10 minutes.

The actual median was 14 minutes 25 seconds. Whoops.

This isn’t an artifact of someone leaving the tab open overnight. Median recorded page time was 14:17, almost exactly the same. Only around 24 % of respondents finished within ten minutes, and roughly two thirds took longer than twelve.

This is even more embarrassing because my own planning document mandated a median above twelve minutes should trigger revision.

Past me had excellent standards for present me. The surprising part is where the time went.

The five actual scenarios, including thirty repeated ratings, took a median of 5:38 combined.

Everything afterward took 8:17.

Post-scenario section Median time
Final attribution questions 2:25
Background / classification 2:11
General beliefs 1:08
Technical knowledge 0:58
Term interpretation 0:47
Demographics 0:29
Summary confidence 0:10

So the experimental design wasn’t really the thing making the survey long.

I had essentially attached another survey to the end of the survey.

This is useful because one of the obvious ways to shorten the instrument would have been to show respondents only a subset of the five scenarios. I actually considered that.

It would also have thrown away one of the nicest properties of the study: every respondent rated every scenario, so the analysis can compare evidence types within the same person.

Given the actual timing data, I’d make the same decision again – I’d just be much more aggressive about everything after it.

Do I really need every general-belief item, can some classification variables disappear?

Could some exploratory questions live in a second study rather than hitching a ride because “we already have respondents here”?

These questions become substantially easier to answer once the median participant has been filling out your questionnaire for fifteen minutes.

People still finished it

The survey was long, but it wasn’t a complete UX disaster.

There were 751 consented sessions in the production data and 538 completions, or 71.6 % completion after consent.

I don’t know how many people saw the landing page and decided they had better things to do with their evening. The stored sessions begin after consent.

Once somebody entered the actual survey though, retention was surprisingly respectable.

740 of 751 reached the first scenario.

608 completed all five scenarios. And of those 608, 538 finished everything else too.

In other words, once somebody survived the actual experiment, they had an 88.5 % chance of enduring the administrative proceedings afterward.

Dropout was gradual rather than concentrated around one page:

751 → 740 → 706 → 679 → 648 → 620 → 607 → 592 → 579 → 565 → 552 → 543 → 539 → 538

There was also no meaningful relationship between which randomized scenario appeared first and eventual completion. Depending on the first scenario, completion ranged from roughly 69.5 to 73.8 %, with essentially no association.

That’s reassuring. The scenarios themselves weren’t obviously scaring people away.

The questionnaire was just long.

The experiment I actually wanted to run

The persistent-agent scenario is the clearest example of something the survey could only approximate.

The fictional system had long-term memory, tools, ongoing tasks and continuity across weeks.

When told that it would be permanently shut down without restoring its memory, it argued for preserving its state and then changed future plans to reduce the chance of losing its memory or tasks.

That gives respondents a vignette about persistence, goals and self-continuity, but it doesn’t give them the experience of interacting with such a system.

The more interesting experiment would have been a long-term one – let someone work with an agent for days or weeks.

Let it accumulate long-term memory around them, remember past conversations, maintain projects, revise plans and build up genuine interaction history.

Then introduce continuity-threatening situations and measure how the person interprets the agent’s response. That would be much closer to the thing I’m actually interested in.

It would also turn my already-too-long fourteen-minute internet survey into a small research program. So I didn’t, for now.

The vignette was a compromise. I still think it was the correct one for this study.

What I’d change

If I built a v2 tomorrow, I wouldn’t redesign everything. The core five-scenario within-subject setup worked well enough that I’d keep it.

I’d also keep the randomized order, the delayed technical-knowledge questions, the neutral survey framing and the post-hoc interpretation checks.

Those decisions were all attempts to protect the primary measurement from the questionnaire itself, and I still think they were right.

I’d want to change three things.

First, Scenario D needs another accessibility pass. Not by weakening the causal evidence, but by finding a simpler way to communicate what was manipulated and what changed.

Second, I’d cut the post-scenario questionnaire. The experimental block was five and a half minutes. There was no good reason for the appendices to take eight.

Third, I’d be more realistic about the time estimate. “7-10 minutes” became a lie despite my genuine attempt not to make it one.

Fourteen minutes is still survivable. I think it’s just polite to tell people.

The slightly annoying conclusion

I don’t think the survey was badly designed; I also don’t think “well, I got 538 responses” magically makes every design decision good.

The instrument was too long by its own standard. One scenario was clearly too technically difficult for a substantial part of the sample.

The two most important conceptual phrases were not interpreted consistently by everyone.

And the experiment I would really like to run couldn’t fit into this format at all.

But the attention check behaved like an attention check. There was no obvious straightlining. Required answers weren’t missing. Most people who completed the core experiment stayed through submission. Randomized scenario order didn’t appear to create a dropout problem.

Most importantly, the things that went wrong are measurable.

I know who failed the technical comprehension check. I know how people interpreted “inner experience.” I know exactly where the time went.

I can run the analysis with and without those groups rather than pretending every respondent understood every sentence exactly as I intended.

A survey doesn’t become valid because its wording looked perfect in the planning document.

It becomes useful when you design enough instrumentation around the questionnaire to discover where actual humans disagreed with your assumptions.

The findings themselves come next.

Back to all findings