Choosing Between Local GenAI and Cloud Models: A UX-Led Decision at a Productivity Startup
AI · 6 min read
The startup's core promise was fast, private drafting inside users' workflows. Early engineering favored cloud models for superior accuracy and easier updates, while the privacy team pushed for on-device inference. UX research found three user segments: privacy-first professionals, speed-sensitive power users, and casual users tolerant of network delays for better output quality.
UX created a decision matrix that mapped user impact across latency, offline capability, data residency, and update cadence. Prototypes compared responses from a lightweight local model and a cloud model with progressive disclosure about source and confidence. Moderated sessions showed users preferred a hybrid strategy that defaulted to local for short tasks and used cloud when users explicitly requested “high-quality” or long-form drafts.
Engineering built a fallback flow: the assistant runs locally and shows a subtle badge with confidence; when confidence is low, a one-tap option sends content to the cloud with an explicit consent modal and a clear estimate of request costs. Early telemetry showed reduced perceived latency and very high opt-in rates for cloud assistance when the benefits were communicated. The UX lesson: surface trade-offs plainly and let users control quality-vs-privacy decisions rather than guessing for them.