Gemini Workspace
2024
Evaluated Gemini AI Integration in Google Workspace for SMBs
Google needed to understand whether Gemini's integration into Workspace would deliver real value for small and medium businesses before broader rollout. Would SMB owners find it helpful enough to adopt? Which features would work best for their workflows?
Understanding if AI features deliver real value to small business owners


Problem area
Could Gemini's AI features deliver real value to small business owners before broader rollout?
Google was preparing to roll out Gemini AI across Workspace tools: Docs, Sheets, Slides, Gmail, and Drive. The team at Google needed to know if small and medium business owners would actually find it useful. Before launching broadly, they needed evidence: which features worked, which fell short, and whether SMB owners would adopt them into their everyday workflows.
When product friction became a retention problem
Problem 2
SMBs lack dedicated IT support, meaning complex or unreliable AI features get abandoned rather than troubleshot
Problem 3
Time is their scarcest resource if a feature doesn't work immediately, they won't invest in learning it
My role
Ensuring data quality and participant management
I supported the multi-session usability study by assisting with data collection, participant coordination and qualitative coding.
My operational role
Data Collection
After each of the 12 tasks, I collected survey responses measuring: Completeness, Understandability, Groundedness and Truthfulness.
Participant Coordination
I managed logistics for 23 participants across multiple sessions, including reminders, troubleshooting technical issues, and quality checks.
Research Approach
How the study was structured
We conducted a mixed-methods study with 23 SMB owners across 3 sessions, testing 12 Gemini tasks through completion rates, surveys, and qualitative feedback.
Demographic details
3 sessions testing 12 tasks total
Mixed-methods: task completion + surveys + qualitative feedback
Key Findings
Reducing errors with smarter confirmations
Where it fell short
Sheets formula creation was the biggest failure 57% couldn't complete the task and 57% rated it "bad," yet it was the most wanted feature with 52% saying they'd use it regularly.
Pattern across all tasks
Simple tasks hit 70-83% completion, but complex multi-file tasks dropped to 35-43%, with over half rating summarization "bad" and 28% flagging truthfulness issues.
Email Drafting
Participants loved getting strong first drafts that captured their tone.
Document Summarization
Bullet-point format made quick comprehension easy.
Sheets Formula
The feature users wanted most didn't work reliably.
What I learned
About research operations
Quality data requires operational rigor. Before this project, I didn't realize how much participant retention impacts research quality. When someone drops out mid-study, you lose the ability to track their experience over time. Maintaining 100% completion meant:
Being responsive to technical issues
Making sessions convenient and well-scheduled
Building rapport so participants stayed engaged
About data quality
Raw data is messy. Survey responses had:
Inconsistent formats (some wrote paragraphs, others single words)
Missing data that needed flagging
Responses that didn't match the questions asked
I developed a systematic cleaning process that balanced preserving participant voice with making data analyzable.
Qualitative and quantitative tell different stories. The numbers showed Sheets formulas had 57% failure rate. But the quotes I pulled revealed why: "I couldn't figure out how to frame the request," "I don't trust the math to be correct." Numbers said "it failed." Words said "users felt confused and lost trust."
About supporting analysis
Good quotes do heavy lifting. The research team's final presentation relied heavily on the quotes I pulled. I learned to select quotes that were:
Representative of common experiences
Specific enough to be credible
Concise enough to be impactful
Context matters when coding. Working with senior researchers to code responses taught me that the same words can mean different things in context. "It's fine" after a successful task means something different than "it's fine" after multiple failed attempts
What I'd do differently
Create a quote database as we go
I pulled quotes after data collection, but organizing them in real-time by theme would have saved time and caught patterns faster.
Build better participant feedback loops
Some technical issues could have been caught earlier with a quick check-in system between sessions.

