Human-AI Interaction
UT Austin
UX Research
Designing Cognitive Friction for Human-AI Collaborative Ideation
A collaborative AI ideation tool that intentionally slows down and adds friction in the process to promote critical thinking and ownership of ideas
AI ideation tools compete on removing effort. But when the thing you are making is an idea, the effort is the work. People take the first polished answer and stop thinking.
Designed the interface and the turn structure for an ideation tool with two kinds of deliberate friction built in. One makes you assemble the idea yourself. The other argues with you. Then we tested both against a plain chatbot.
Fast, automatic thinking fell from 79% to 42%. Four of five people quit the plain chatbot before it was finished. Only one of the two friction mechanisms also made people feel the idea was theirs.
The Problem
























The ideas that lose to filing friction are the ones nobody else has.
We audited what people actually reach for when they make things. Every one of them competes on the same axis: less effort between wanting a thing and having it. That promise is a good one when you already know what you want. Book the flight. Resize the image. Write the email you have already composed in your head. The faster the tool gets out of the way, the better it is.
An idea is not like that. You do not know what you want yet. Working it out is the task, and the thinking is what does the working out. So when a tool removes the effort, it has not saved you time. It has removed the only part that was getting you anywhere.
People take what the model hands them and move on. The generative half of the work quietly moves to the machine.
Person and model converge within a turn or two. There is no disagreement left to explore.
An idea that arrives finished is easy to abandon. Effort is what makes it yours.
Why it happens
A polished answer never trips the wire
Dual-process theory splits cognition in two. System 1 is fast, associative, heuristic and effortless. System 2 is slow, procedural, metacognitive and effortful. Creative work needs both, because System 1 generates and System 2 evaluates.
The System 1 to System 2 transition is not voluntary. System 2 engages when System 1 hits an error it cannot resolve, so confusion is the trigger. AI-assisted ideation is documented as System 1 dominant, and this is why: fluent output produces no error to catch.
The fluency trap: output that is syntactically perfect and conceptually shallow. Nothing looks wrong, so nothing triggers System 2.
Research
Seven kinds of friction exist.
We could build two.
Deliberate friction has a literature. Under desirable difficulties in educational psychology, and positive friction and frictional AI in interaction design, we found seven families.
Two questions cut it down. Can three students build it in one semester, and does it change what the person has to do rather than just how the screen looks. Perceptual disfluency and temporal barriers fail the second test. They decorate the problem instead of touching it.
Degraded fonts, hard-to-read typesetting. Make it harder to read.
Microboundaries, artificial processing delays. Make the person wait.
Effortful assembly, the IKEA effect. Make the person do the work.
Forced choice, unassisted first steps. Withhold help until they commit.
Engineered discomfort, unfriendly interfaces. Make the system argue.
Scaffolding, guiding questions. Give hints instead of answers.
Reflective nudges, explainability requirements. Make someone check the reasoning.
Generative labour and agonistic design were the only two that were both buildable and interactional.

Project Ideas
How to make an experiment around them
Degraded fonts, hard-to-read typesetting. Make it harder to read.
Microboundaries, artificial processing delays. Make the person wait.
Forced choice, unassisted first steps. Withhold help until they commit.
Scaffolding, guiding questions. Give hints instead of answers.
Reflective nudges, explainability requirements. Make someone reason.
Four of these test what you already know. Only the last is about making something, and it is the only one that shipped. It became the Antagonistic Partner.
Response
So we built three tools
We built a text ideation tool for marketing briefs. It has three modes behind a single toggle. The model, the task and the length of the session are identical in all three. The only thing that changes is how much work the interface makes you do.
You describe a brief, it hands you finished campaign ideas, you pick one. No canvas, no resistance. This is the control, and it is what every tool on the market already does.
The model puts rough fragments on a canvas and stops. You cannot advance a turn until you have edited them, added detail and drawn the connections yourself.
It is helpful for two turns. Once you commit to an idea it turns critic, acknowledges without praise, and pushes on your weakest answer, harder every turn.
Research Questions
1
Does generative labour or an antagonistic conversation style trigger System 1 to System 2 transitions during ideation?
2
Do they affect psychological ownership of the resulting idea, and the contribution the person attributes to the AI?
Nobody was told which mode they were in, or that the mode changed between sessions. The next four sections are how each mode works. After that, what happened when people used them.
The control condition. A plain chat partner, no canvas, no resistance, and what every tool on the market already does.
If the two friction modes differed in more than one way, no result could say which difference did the work.
Everything except the friction mechanism was held identical. Same model, same brief format, same turn cap, same interface shell.
The control condition. A plain chat partner, no canvas, no resistance, and what every tool on the market already does.

Model 1
The model opens the space. You make the meaning.
Chat on the left, a node canvas on the right. Every turn the model puts something abstract on the canvas, and the session will not move until you have worked on it. Edit the nodes, add detail, draw the connections yourself.

Turn four of six. Directions at the top, channels in the middle, named concepts at the bottom, all wired by the person using it.
1
Three questions. Nothing on the canvas yet.
2
Directions arrive as single words. You pick.
3
Channels wire to the direction you chose.
4
Named concepts built from your choices.
5
Three empty fields. The model writes nothing.
6
It reads back the idea you built.
One turn does the real work
Turn five is the only turn where the model writes nothing. What will make the audience act, what will they do, and what happens next. Every turn before it was choosing. This one is writing, and it is where people stalled.
Effort that feels like busywork produces resentment, not thinking. An open canvas where you connect anything to anything made the labour feel arbitrary.
Six fixed turns walking abstract to concrete, so every action builds on the one before it and the effort has somewhere to go.
The cost: the session is rigid. A confident user walks the same six steps as everyone else.
Friction needs an exit, so we built four and scoped each one tightly. Undo on the canvas. Regenerate, but only on turns one to four, while ideas are still cheap. Go Back, restoring the canvas as it stood at turn four. And a hint button that exists on turn five and nowhere else, because that is the only turn people actually stalled on.
Model 2
It is friendly for two turns, then it turns
It was going to be a shape-shifter. The plan had the model role-playing a different stakeholder each turn, a marketer then a customer then a sceptic. That was cut for one critic who escalates, because a rotating cast gives you three shallow objections instead of one that gets deeper.
For the first two turns it is a neutral evaluator. It clarifies your brief, hands you three or four named campaign ideas, and asks which one excites you. Warm, useful, ordinary. The moment you choose, the interface swaps the prompt underneath you. From turn three it is a critic. It acknowledges your idea without praising it, asks questions built to find the weak joint, and never fully accepts an answer. Each turn it pushes harder.


Two personas behind one mode. The swap happens in the front end, at turn three, the moment after you commit to an idea.
Hostility at turn one has nothing to bite. You have not committed to anything yet, so pressure reads as noise and people disengage.
The persona switch waits until after the choice. Every challenge then lands on something the person has already claimed as theirs.
The cost: those first two friendly turns are a small ambush. People trusted it, then it turned on them.
One pilot participant said the antagonistic mode made them feel they were defending someone else’s idea. That showed up in the ownership scores months later. It is the honest limit of this mechanism.
Method
Most of this interface is a string
The canvas is the part you see. The part that decides how a turn behaves is a prompt. Each of the six turns carries its own instruction set, and every reply comes back as strict JSON against a fixed contract, so the interface can place nodes without guessing at them.
The canvas state is fed back to the model on every single message. It reads what is already there, returns only what it is adding, and never re-lists the rest. That is what stops six turns of accumulated work being quietly rewritten.
React and ReactFlow for the canvas, Zustand for state, a thin Express server, and llama3 served locally through Ollama. The interface was drawn in Figma first, then built out with Claude Code, by the three of us together on one machine.
Running the model on a laptop was a constraint that shaped the design. No session ever left the room, which is what let us record everything. But a small local model cannot be trusted to freewheel, and that is the real reason each turn carries its own instruction set and a strict output contract rather than one open-ended prompt. 2,056 lines across the tracked source, versioned in git. Four commits, because we worked in one room on one screen rather than in branches.
A study about ownership is worthless if the model can edit the thing the person made.
Every node a user creates carries an id the model is forbidden to modify, enforced twice: once in the prompt, once in the code that handles the reply.
The cost: the model sometimes refers to user work clumsily, because it cannot tidy it. Correct trade for this project.
Five orders across five people. Every mode appears in every position exactly once. Whichever mode someone tries first gets their freshest attention, and whichever comes last gets a tired participant, so we made a counterbalanced Latin square. Five orders across five people so position cancels out.
P1
P2
P3
P4
Analysis
How words become data
We wrote the codebook before anyone looked at a transcript. Three top level codes, twelve underneath, built from dual process theory rather than from whatever we happened to notice.
System 1
System 2
Cross-cutting
inter-rater reliability on the top-level System 1 / System 2 codes. Substantial agreement.
inter-rater reliability on the twelve sub-codes.
Weaker, and expected.
The cost: the model sometimes refers to user work clumsily, because it cannot tidy it..
Two of us coded independently and reliability was calculated as Cohen’s κ. Coders agreed strongly on whether System 2 was engaged and less strongly on which sub-process it was, which is the expected pattern for theory-driven qualitative coding.
The third category is the interesting one. A person starting on instinct, then catching themselves. Those moments were rare, 3.65% of all codes, and they appeared only in the friction modes, never in the baseline.
Results
Both mechanisms changed how people thought. Only one changed how they felt.
Everything below comes from five participants running fifteen briefs across three modes, with every session recorded and replayed turn by turn. Here is the whole picture on one line each.

In the baseline, four of five people stopped before the tool did, and two quit after three turns. Asked why, they said they had “no more to discuss”. That is premature convergence, and it is the failure the whole project was built to interrupt.


Psychological ownership 3.4 to 5.2. Credit given to the AI fell by two points. Confidence went up. Satisfaction did not move. It won on every measure the study had.
Psychological ownership 3.4 to 3.6. Statistically nothing, on five people. It made people think harder, then cost them 1.4 points of satisfaction and 0.6 of confidence.
System 2 engagement and psychological ownership are separate outcomes, and only one of the two mechanisms moved both.
Findings
What we found
Both mechanisms roughly doubled engagement, so the summary numbers make them look like two versions of the same idea. The timing logs say otherwise.
Psychological ownership 3.4 to 5.2. Credit given to the AI fell by two points. Confidence went up. Satisfaction did not move. It won on every measure the study had.
Psychological ownership 3.4 to 3.6. Statistically nothing, on five people. It made people think harder, then cost them 1.4 points of satisfaction and 0.6 of confidence.
System 2 engagement and psychological ownership are separate outcomes, and only one of the two mechanisms moved both.
