Granite Playground

Live
IBM's open-source enterprise AI models are called Granite. Granite Playground lets people try chat, search, and research-agent capabilities directly in the browser.

2.6× daily prompts · 3× unique users vs Granite 3.3

Overview

Role & Responsibilities
Hands-on design lead across end-to-end product design (User research, product strategy, UX design, interaction design, and prototyping)
Period
2025 Q2
Platform
Web
Team
Designer × 3, Engineer × 5, PM × 1

What I did

  • Defined user flows for the targeted users and enhanced their experience
  • Owned the design direction for the landing page and sample prompts experience
  • Guided interaction patterns and microinteraction quality across the product

Brief

With Granite 4.0, IBM wanted to drive broader adoption among enterprise customers. But while competitors had made their models effortless to try, Granite still sat behind setup and documentation. The launch needed a more accessible way in.

So the Playground shipped day 0, alongside the model: the first time anyone could use an IBM model without an IBM account, built to serve two very different audiences at once.

Users

Enterprise Developers

Potential adopters who need to test Granite with their own use cases and understand the model's capabilities.

IBM Client Managers

Non-technical stakeholders who need relevant examples they can tailor and demonstrate directly to customers.

Design challenge

How might we give enterprise developers and IBM client managers an easy way to try Granite while supporting their different goals?

Process

01

User Research

Mapping two journeys

Before designing screens, I mapped the end-to-end journey for both user groups to identify friction and design opportunities. Enterprise developers wanted to test Granite with their own use cases and understand the model's capabilities. IBM client managers needed examples they could tailor to the customers they were presenting to.

One shared entry point

I proposed a curated, filterable sample-prompt library that gave both groups one place to start while still helping each find what they needed. The experience also had to work within anonymous sessions and per-user token budgets without cutting people off mid-evaluation.

Problem definition

How do we serve both technical evaluators and business decision-makers in a single experience?

Key insights
  • Developers measure accuracy; managers measure business impact
  • Generic prompts create skepticism; industry-specific ones build instant credibility
Alternatives considered
  • Separate role-based onboarding paths
  • Community-contributed prompts
Design rationale

A unified curated experience with enterprise-focused prompts, technical depth with business relevance.

02

Sample Prompts: The Front Door

A useful place to start

I designed a sample-prompt library for users facing a blank prompt box. Industry hashtags made relevant examples easy to find and tailor for specific customers.

Managed in Airtable

Instead of building an admin page, I used Airtable to store and serve prompts. Non-technical team members could keep the library current without hard-coding changes or redeploying the service.

Enterprise-focused prompts

Every prompt reads like real enterprise work, not a generic AI demo, so the first result a visitor sees looks like their own work.

One click, full context

Cards show a short summary, then reveal the full prompt on hover. Users keep context while submitting in one click.

Airtable instead of an admin page

Non-technical team members can add, edit, and recategorize prompts without hard-coding or redeploying.

Problem definition

What do first-time users do when they don't know what to ask an enterprise model?

Key insights
  • Generic prompts create skepticism; industry-specific ones build instant credibility
  • Users scan rather than read; a prompt card has to communicate its value before it's clicked
Alternatives considered
  • A blank chat with placeholder suggestions
  • Separate role-based onboarding paths
Design rationale

One curated library serves both audiences: industry hashtags help users find relevant prompts, and editable examples let client managers tailor them for customers.

03

A Chat Interface Built for Trust

Show the work

Enterprise evaluation lives or dies on trust, so the chat interface was designed to show its work: multi-turn conversations with visible source citations, a research agent that streams its trajectory live (its plan and every source it adds, as it adds them), and report generation that turns a session into a shareable document.

Transparency over magic

Users see what the model is doing, where information came from, and what happens next. For developers that's evidence; for client managers it's a demo that holds up in front of a client.

Trajectory, not a spinner

The research agent streams each source as it finds it, so a long run reads as work you can watch instead of a loading state.

Citations as evidence

Every answer links back to where it came from, turning output into evidence users can verify rather than text to take on faith.

Problem definition

How do we serve both technical and non-technical users in one interface?

Key insights
  • Transparency in AI interactions builds trust. Users want to understand what's happening
  • Users scan rather than read. Key information must be immediately visible and scannable
Alternatives considered
  • Split-screen design with live preview and configuration panels
  • Conversational UI prioritizing natural language interaction over structured inputs
Design rationale

Card-based layout with clear visual hierarchy, supports progressive disclosure, familiar patterns, and scales across devices.

04

Design System & Development Handoff

Design that survives implementation

Handed off the playground as a complete design system: a component library, breakpoint specs for every layout, and annotated interaction states, so engineers could build without guessing at intent.

Staying through QA

Handoff wasn't the end of the design work. I stayed embedded through implementation with collaborative QA passes, adapting the design as technical constraints surfaced and refining details against real model behavior, which is why the shipped product matches the design.

Impact

Granite Playground became the primary activation surface for Granite 4.0, driving 2.6× more prompts per day and 3× more unique users than Granite 3.3, and directing roughly 30% of total Granite web traffic to the Playground and Docs.

These outcomes contributed to IBM's most successful Granite launch to date across Day 0 metrics, shifting Granite from something developers had to evaluate into something they could experience.

Learnings

  • How to design enterprise-focused playgrounds that address both technical and non-technical user needs.
  • The importance of mapping user flows to identify the pain points and the user needs. Based on this analysis, we decided to have the sample prompts tailored to real-world enterprise use cases.
  • Clear, transparent UX helps demystify AI model capabilities for diverse audiences.
  • How to increase the fidelity of the design to the level of public-facing product

Next steps

  • Iterate on the design based on the usage analytics
  • Add more agents to the playground