Sitemap

Bridging the Usability Gap in LLM Tools

A UX Research & Design Case Study

--

Press enter or click to view image in full size
Favorite Feature: A star icon to save a response and return to it later.

Team: 4 Researchers & Designers
My Role: UX Researcher & Interaction Designer

I led the research framing and problem definition, contributed across all four prototyping stages (10×10 ideation, storyboarding, paper prototype, high-fidelity Figma), and independently facilitated think-aloud sessions with my four participants.

The Problem Nobody Was Naming Correctly

When we started this project, the obvious framing was: AI hallucinates. Let’s help users catch that.

It took one exercise to realize how wrong that framing was.

The real problem isn’t that AI gets things wrong. It’s that users have no way to know when it does. The response arrives fluent, confident, and well-formatted. There are no error signals, no confidence levels, and no sources surfaced by default. The interface is designed to feel authoritative, and that’s exactly what makes it dangerous.

We were designing for ChatGPT users (students, researchers, early-career professionals) who were already relying on the tool for consequential work. Our preliminary research pointed to three interconnected gaps that nobody was directly addressing:

Prompt Quality Gap. Users, especially students writing research papers, don’t know how to phrase prompts to get reliable, well-sourced responses. Vague prompts produce vague outputs, and the user doesn’t know what they missed.

Trust & Verifiability Gap. There’s no in-interface mechanism to evaluate how confident or well-sourced a response is. Users can’t distinguish accurate output from hallucinated output, and most don’t try to verify.

Cognitive Load & Retention Gap. Responses are long, complex, and ephemeral. Users have no quick way to save what matters or get a simpler version when they need one.

The design challenge wasn’t “detect hallucinations.” It was: design trust across the full interaction lifecycle.

Generating Ideas: The 10×10 Exercise

We ran a 10×10 sketching session, ten concepts in ten minutes, to break out of obvious solutions fast.

Press enter or click to view image in full size
Press enter or click to view image in full size
Press enter or click to view image in full size
Press enter or click to view image in full size

Early sketches gravitated toward post-hoc error flags: pop-ups that appeared after a response, warning labels on suspicious sentences. These felt logical at first. But as we kept sketching, a pattern emerged: every intervention that happened after the response was already too late. The user had already read it, already half-trusted it.

The exercise forced a reframe. By sketch seven or eight, we were drawing interventions at three distinct moments:

  • Before the question, helping users write prompts that would produce better responses
  • During the response, making the AI’s confidence visible inline
  • After reading, helping users save, simplify, and verify what they’d received

That three-point model became the organizing logic for everything that followed.

Storyboards: Seeing the User, Not Just the Interface

We built two storyboards: one showing the current problem and one showing the proposed solution.

Press enter or click to view image in full size

The current-state storyboard followed a student writing a history report. She asks ChatGPT about the French Revolution. The response is fluent and confident. She uses it. Her report includes a hallucinated date and a fabricated causal claim. She doesn’t find out until her grade comes back.

What the storyboard made visible: she wasn’t negligent. She had no signal to act on. The interface gave her no reason to doubt.

The solution storyboard showed the same student, same question, but now she sees a prompt improvement suggestion before she sends, a confidence score and source list with the response, a simplify toggle for the dense paragraph, and a favorite button to save the parts she trusts. By the end, she’s using AI as a tool she understands rather than an authority she defers to.

The storyboard phase shifted our language internally. We stopped saying “reduce hallucinations” and started saying “give users the signals they need to verify.” That distinction shaped every design decision afterward.

Paper Prototype: Where the Obvious Solutions Broke

Press enter or click to view image in full size

Paper prototyping exposed problems that had looked fine in sketches.

The confidence score was too abstract. Our first version showed a percentage, “74% confidence,” next to the response. Users in walkthroughs stared at it without knowing what to do. We added a color-coded label (Low / Medium / High) so the meaning landed at a glance without requiring interpretation.

The prompt improvement step felt invasive. On paper, showing a rewritten version of the user’s prompt before they sent it felt like the system was correcting them. We changed it to a soft suggestion: “Here’s an improved version. Use it, edit it, or skip it.” That single framing shift changed it from an interruption to a helpful offer.

The analytics view was cut entirely. We had sketched a dashboard showing the user’s AI usage patterns and confidence history. On paper it was immediately clear this was feature creep; it pulled attention away from the core moment of trust. We cut it and narrowed the flow to six core screens.

The lesson from paper: fewer screens felt more trustworthy. When users could take in everything on the screen, they felt in control. When they couldn’t, the complexity itself became a trust signal, a negative one.

High-Fidelity Prototype: What Only Digital Could Answer

Moving to Figma surfaced a different class of questions. Paper couldn’t tell us whether the confidence indicator felt credible or gimmicky, that’s a visual hierarchy and typography question. Digital could.

A star icon to save a response and return to it later.
Press enter or click to view image in full size
Press enter or click to view image in full size
Press enter or click to view image in full size
Screens 1 & 2: Simplify mode automatically simplifies responses into shorter, clearer, easy-to-read explanations. Screen 3: Evaluates how reliable each statement is based on supporting evidence and sources. Screen 4: Improve the prompt with more detail and a clearer question for better responses.

Key decisions at this stage:

Source pills needed click affordances. On paper, source references looked fine as static labels. In the digital prototype, users didn’t know they were interactive. We added subtle underlines and hover states so it was clear they could explore further.

Visual weight of the confidence score mattered more than we expected. Too prominent and it felt alarming; too subtle and users didn’t notice it. We landed on a mid-weight treatment, present and readable, but not the first thing demanding attention.

The prompt improvement step needed a very light visual tone. Any suggestion UI that looks like a warning creates anxiety. We used a soft card with an edit field and gentle copy (“This prompt could be improved”) rather than anything that read as a correction or alert.

Think-Aloud Testing: What Users Actually Did

As a team, we conducted formative think-aloud sessions with 16 participants across four researchers. I independently recruited and ran four of those sessions, each with a different participant, to surface design direction before a larger validation study.

What worked well:

Across all sessions, every core feature received positive responses, no participant rejected or ignored any of the four features. Specific observations:

  • Participants found the overall flow easy to navigate without instruction
  • The color-coded confidence labels (Low / Medium / High) were immediately understood and appreciated
  • Most participants accepted the rewritten prompt suggestion without editing, they trusted it, which validated the soft framing
  • Several said the confidence score made them feel more confident in the AI’s output, which was exactly the calibration effect we were designing for

What didn’t work:

  • Participants couldn’t distinguish between “Helpful” and “Favorite”, two actions we’d treated as meaningfully different but which looked and felt nearly identical. This was a clear information architecture problem, not just a labeling one.
  • Some participants asked for keyboard shortcuts, suggesting the power-user workflow wasn’t fully addressed
  • The source verification step was noticed but not always acted on, users trusted the cues more than they verified the underlying sources, which raised a follow-up question: are confidence indicators actually reducing verification behavior by making users feel they don’t need to check?

That last finding was the most interesting tension to emerge from testing. We’d designed signals to help users evaluate trust. But some users were treating the signals as the verification, rather than as a prompt to go verify. The design was working, and also potentially creating a new form of over-reliance.

How the Design Changed

Press enter or click to view image in full size

What This Project Clarified

The clearest lesson wasn’t about any single design decision, it was about the process.

Each stage of prototyping answered a different question. Sketches asked what could this be? Storyboards asked why does it matter? Paper asked does the flow work? Digital asked does it feel right? Compressing or skipping any of those stages would have produced a worse outcome, not a faster one.

The constraint that helped most was limiting the paper prototype to a few screens. It forced us to cut features that were genuinely interesting but would have diffused the core idea. Trust design doesn’t benefit from more features, it benefits from fewer, clearer signals.

And the tension we didn’t fully resolve, whether confidence indicators help users verify or give them permission not to, is the most interesting question this project opened. Users don’t just need accurate AI outputs. They need signals that help them evaluate, interpret, and trust those outputs appropriately. Designing for appropriate trust, rather than just more trust, is harder. It’s also where the most important design work in this space still needs to happen.

--

--