Gemini Live

2024

Shaped Gemini Live's product roadmap through diary study

Weather app image
Weather app image
Weather app image
Weather app image

Role

Junior Researcher

Team

2x Project Managers

2x Researchers

1x Recruitment Team

Impact

Identified a trust gap where users anthropomorphized Gemini early but hit a wall when it lost context mid-conversation

Mapped drop-off triggers to moments where Gemini gave shallow responses to follow-up and high-effort tasks

Informed feature prioritization by pinpointing generic responses in high-stakes interactions as the primary driver of user drop-off

Role

Junior Researcher

Team

2x Project Managers

2x Researchers

1x Recruitment Team

Impact

Identified a trust gap where users anthropomorphized Gemini early but hit a wall when it lost context mid-conversation

Mapped drop-off triggers to moments where Gemini gave shallow responses to follow-up and high-effort tasks

Informed feature prioritization by pinpointing generic responses in high-stakes interactions as the primary driver of user drop-off

Problem area

Initial launch of Gemini Live

Google was preparing to launch Gemini Live, a new conversational voice AI, and needed to understand how real users experienced it in everyday life. I joined a cross-functional team to run a 2-week unmoderated diary study with 34 participants, capturing how people naturally integrated Gemini Live into entertainment, learning, and personal conversations.

Research Goals

What the team needed to learn

Use Case Validation

Understand how users integrated voice AI into daily workflows

Use Case Validation

Understand how users integrated voice AI into daily workflows

First Run and Long Term Experience

How users' perception of Gemini evolved over 2 weeks

Conversation Flow

Assess how well Gemini Live supports complex conversation topics including acting as a coach and navigating personal conversations with users

My role

What I focused on

Thematic Analysis

Identified key themes across 4 use case categories

Thematic Analysis

Identified key themes across 4 use case categories

Coding Data

Applied coding frameworks to transcripts and session notes to identify patterns

Findings Documentation

Created findings presentation and strategic recommendations for product team

Participant Management

Provided one-on-one support to resolve questions about study instructions

Use Case 01: Activating gemini live

How did users react to setting up Gemini Live

Small UI Details Created Big Friction
While reviewing activation videos, I noticed participants repeatedly struggling with the same interface elements. I coded these friction points:

Challenges Identified

Icon Visibility Challenge

The "Live" button lacked distinctiveness, participants suggested color gradients, motion effects, or more prominent visual cues to make it stand out from other interface elements.

Status Uncertainty

Users couldn't confidently tell when conversations were paused, often seeing audio waves move even when supposedly on hold, creating confusion about system state.

Quality of voices

Some participants found the voices too robotic or not appropriately toned, with some describing them as too sassy or too friendly.

Participant Quotes

Icon Visibility Challenge

"I think the small button may not stand out to users who don't have prior instructions"

Status Uncertainty

"I can't tell with hundred percent certainty that the hold button is working, I see the waves moving, and so maybe it is listening to me.."

Quality of voices

Some of the voices sounded a bit too robotic and not natural-sounding.

Use Case 02: Conversations about entertainment topics- books, movies, TV

How can Gemini Live hold natural entertainment conversations without losing context?

After analyzing participant conversation transcripts across books, movies, and TV shows to understand engagement patterns, we found that:

  • Books: Gemini performed best in conversations about books. Detailed insights without spoilers kept users engaged.

  • TV Shows: Its responses regarding TV shows were Inconsistent: Newer and niche content led to factual errors, due to training data lag.

  • Context Loss: Across all entertainment topics, Gemini routinely lost conversation threads during follow-ups, breaking the flow of discussion.

Transcript of Gemini losing context mid-conversation:

Challenges Identified

Popular v Niche Content Divide

Books generated the most satisfying conversations, while TV show discussions were inconsistent, especially with less popular content leading to factual errors.

Context Loss Issues

Gemini lost conversational thread during follow-up questions, such as confusing "Mercedes" the character with Mercedes the car brand during a Count of Monte Cristo discussion.

Missing Visual Cues

Without visual cues like covers or images, participants struggled to stay engaged and wanted links to streaming platforms or purchase options to deepen the experience.

Participant Quotes

Popular v Niche Content Divide

"Gemini said many contradicting things. Initially it said it hasn't heard about Young Sheldon. After a lot of back and forth, it was able to look it up but my questions were incorrectly answered."

Context Loss Issues

"The conversation got derailed in the middle where Gemini didn't remember we were talking about The Office which wasn't ideal

Missing Visual Cues

"I would have wanted for Gemini to show me in the background where to watch the show and an image of the TV show."

Use Case 03: Acting as a coach

How well does Gemini Live support users learning new skills?

Through coding learning conversations (piano, coding, ukulele), I discovered a mismatch between user expectations and Gemini's output.

Challenges Identified

Skill Level Mismatch

Asking for skill level would make the experience more valuable. Advanced learners were underwhelmed and beginners sometimes found the guidance too complex. 

User Expectation Mismatch

Participants expected Gemini to complete learning tasks (like creating resumes, setting up practice schedules) rather than just providing instructions.


Need for Visual Aids

Voice-only explanations became awkward for visual skills like coding or ukulele playing, participants needed demonstrations, not just verbal instructions.

Participant Quotes

Skill level mismatch

"This one was a little more difficult just because I was asking for tips on learning a skill that is typically learned visually and hard for Gemini to explain verbally."

User Expectation Mismacth

"It was a little awkward to have Gemini Live explain coding syntax through conversation as you can't exactly communicate new lines and such."

Need for visual aids

"It can ask us to pick a skill level and give us response accordingly. This can be a global setting or per skill I want to learn."

Use Case 04: Personal conversations with Gemini

How can Gemini Live create natural dialogue about personal challenges?

Through coding learning conversations (piano, coding, ukulele), I discovered a mismatch between user expectations and Gemini's output.

Challenges Identified

Grasp of nuance

Participants valued Gemini's ability to understand and address the subtleties of their queries.

Adequate understanding on nuance

Participants valued Gemini's ability to understand and address the subtleties of their queries.

Lack of personalised responses

Gemini’s responses were not tailored in a helpful way. Many responses were very standard, such as “be more honest,” which were seen as generic pacifiers and not very useful.

Issues with turn-taking

The turn-taking needed improvement for difficult conversations to be well served. Gemini did not understand the user’s context, resulting in more generic responses.

Participant Quotes

Grasp of nuance

"I really liked the phrasing Gemini used on reaching out to a friend that owes me money. It is a difficult subject to bring up but it made it sound so easy."

Lack of personliased responses

"The conversation felt that it was definitely attempting to be helpful, but in terms of actionable suggestions, I don't think there was much that I wouldn't already think of on my own."

Issues with turn-taking

"Gemini should ask for more context to inspire deeper thinking/finding the most appropriate solution."

Outcomes and Retrospective

What I learned about AI and expectations

This project taught me that unmoderated diary studies require reading implicit signals like tone changes, repeated attempts, and abandon points, not just explicit feedback. I also learned that people anthropomorphize conversational AI quickly, treating Gemini like a friend within days, which meant generic responses felt like personal rejection. My most valuable contribution wasn't documenting what worked but mapping the gap between what users expected, task completion, and what Gemini delivered, instructions. That delta became the roadmap. Working on evolving technology also forced me to distinguish between product bugs that needed immediate fixes and fundamental design questions about what the product should be.

Outcomes

Shipped example prompts

Research directly informed product implementation. Gemini Live added on-screen example prompts based on participant feedback requesting guidance for brainstorming, skill-learning, and coaching tasks

Shipped example prompts

Research directly informed product implementation. Gemini Live added on-screen example prompts based on participant feedback

Shipped example prompts

Research directly informed product implementation. Gemini Live added on-screen example prompts based on participant feedback requesting guidance for brainstorming, skill-learning, and coaching tasks

Secured follow-on work

Documented 15+ additional feature requests including playback speed controls and source attribution that led to team referral to YouTube Music for similar research

Secured follow-on work

Documented 15+ additional feature requests including playback speed controls and source attribution that led to team referral to YouTube Music for similar research

Identified trust-breaking patterns

Surfaced instances of Live claiming human experiences ("I watched that movie") that undermined conversational authenticity

Identified trust-breaking patterns

Surfaced instances of Live claiming human experiences ("I watched that movie") that undermined conversational authenticity

Retrospective

More structured comparison framework

I coded conversations organically, which worked, but a pre-defined rubric for "conversation quality" would have made cross-participant comparisons more systematic.

More structured comparison framework

I coded conversations organically, but a pre-defined rubric for conversation quality would have made cross-participant comparisons easier.

More structured comparison framework

I coded conversations organically, which worked, but a pre-defined rubric for "conversation quality" would have made cross-participant comparisons more systematic.

Quantify patterns earlier

I identified themes qualitatively first, then went back to count occurrences. Next time I'd track frequencies during initial coding to spot patterns faster.

Quantify patterns earlier

I identified themes qualitatively first, then went back to count occurrences. Next time I'd track frequencies during initial coding to spot patterns faster.

User journey mapping

I could have created visual maps of where conversations broke down to make findings more immediately actionable for designers.

User journey mapping

I could have created visual maps of where conversations broke down to make findings more immediately actionable for designers.

Create a free website with Framer, the website builder loved by startups, designers and agencies.