RWC 2. (~6 min)

Early AI exploration

This case illustrates how I approach unprecedented challenges, adopting emerging technologies and navigating highly ambiguous problem spaces. To respect NDA constraints, high-fidelity visuals have been replaced with simplified, low-fidelity representations.

Mission

“Investigate how emerging LLM technology could create tangible value and measurable efficiency gains for support reps.”

Context

This was one of the company’s first ventures into AI technology. As LLM capabilities began to emerge, the organization decided to explore how this new technology could be integrated into their connected ERP software system. The Case Management area was chosen as the initial focus, a space where support reps spend significant time trying to understand and resolve customer issues. The challenge was simple yet profound: could we leverage LLMs to genuinely help our users, or was this just another tech trend that wouldn’t deliver real value?

Problem

  • Define meaningful use cases for LLM technology in the absence of clear precedents or validated user needs.
  • Balance exploration with accountability, ensuring experiments translated into measurable user value rather than novelty.
  • Operate under technical and organizational uncertainty while establishing patterns that could scale.

Constraints

  • Scope: Limited to Case Management area only for this initial AI exploration.
  • Organizational: No established AI governance or adoption policies, creating uncertainty in decision-making.
  • Industry precedent: Little to no reference for LLM integration in ERP systems—we were pioneering new ground.
  • Technical integration: Had to work within the live production system without disruption.
  • Technical limitations: Early LLM implementations only supported single-turn responses—no conversational capability to refine or clarify output.
  • Model availability: Started with only Cohere and Llama; ChatGPT access came later.

My role

  • Technology translation: Understood LLM capabilities and bridged the gap between technical possibilities and user needs.
  • Strategic planning: Facilitated workshops, created roadmaps, and coordinated cross-functional research efforts.
  • Cross-functional alignment: Collaborated with PM, developers, and researchers to ensure cohesive execution.
  • Quality advocacy: Championed the creation of evaluation frameworks to measure and improve AI output quality.
  • Resourceful testing: Advocated for and conducted user testing with SMEs when dedicated research resources weren’t available.
  • Role definition: Helped establish what UX design means in the context of AI-powered products.

Exploration

1. Opportunity Discovery

  • Ran cross-functional workshops to identify high-impact use cases within Case Management aligned with LLM capabilities.

  • Mapped ideas by user value vs. feasibility, resulting in three priority areas:

    1. Case Summary: Strong alignment with long-term product vision. Addresses a major user pain point by reducing time to comprehension.

    2. Similar Cases: Low-complexity, high-value “quick win.” Improves discoverability of relevant historical knowledge.

    3. Proposed Solutions: Deferred. Required more mature models and specialized agents to ensure reliability and trust.

Opportunity map comparing potential AI use cases by user value and feasibility.

2. Case Summary

Goal: Reduce cognitive load and accelerate understanding of complex cases.

Approach

  • Conducted SME interviews to identify the critical information needed for case comprehension.
  • Enabled an internal prompt-testing environment using real (anonymized) data.
  • Logged prompts, outputs, and improvement hypotheses in a shared repository.
  • Developed a UX evaluation framework to compare generated summaries against “ideal” expert outputs.
  • Iterated on prompt structure and tone based on qualitative and quantitative feedback.
  • Validated results through repeated SME review cycles.

3. Similar Cases

Goal: Help agents quickly identify relevant past cases and reuse knowledge.

Approach

  • Interviewed SMEs to define similarity criteria (context, issue type, resolution patterns).
  • Explored ranking and confidence indicators to build user trust and transparency.
  • Tested presentation formats and relevance scoring with SMEs and pilot users.
  • Refined labeling and explanation patterns to reduce ambiguity.

4. AI Interaction Standards

Goal: Ensure consistency, scalability, and quality across AI-generated outputs.

  • Established shared guidelines for tone, structure, and terminology in summaries and insights.
  • Defined reusable prompt patterns and formatting rules to maintain coherence.
  • Created internal documentation to support cross-team adoption and iteration.
  • Focused on repeatability and consistency rather than one-off prompt tuning.

Decision points

1. Assessing AI Output in Real Scenarios

The goal from the start was to be realistic about output quality, not optimistic. Rather than judging the system on best-case demos, we wanted to understand how it behaved across real cases.

To do that, we enabled a workflow that allowed us to quickly iterate on prompts and test them against different case histories. Each output was then compared to an “ideal” summary for that situation, defined with support from SMEs.

I evaluated results using Nielsen–Norman’s 10 usability heuristics, which gave us a consistent way to assess quality, track progress, and identify where the system needed improvement.

Evaluation workflow comparing generated case summaries against expected expert output.

2. Summary: overview + timeline structure

Through testing, I discovered the LLM struggled to synthesize the full context of complex cases, leading to incomplete or misleading summaries. To build trust and reduce errors, I designed a hybrid output: a concise two-sentence overview followed by a chronological timeline where each activity was individually summarized. This solution aligned perfectly with our broader vision, while directly addressing the core pain point: support agents were drowning in scattered information that took too long to piece together and understand

Case summary structured as a concise overview followed by a chronological timeline.

3. Similar Cases: Helping Users Judge Relevance

LLMs operate as black boxes, while we could access similarity vectors and convert them to percentages, these technical metrics didn’t align with how support agents actually think about case relevance. After deliberation, I decided to hide the percentage scores and simply present similar cases in ranked order. To help users judge usefulness themselves, I displayed only what mattered: the case title and a summary of the initial customer message, giving reps the context they needed to decide if a case was worth exploring further.

Similar cases presented in ranked order with titles and customer-message summaries.

4. Establishing a Pattern: Meaningful Waits, Smart Costs

This project launched during the earliest days of LLM adoption in the company, no established patterns existed, and we had the opportunity to set precedents. Leadership grew increasingly nervous about the per-call costs of LLM requests, creating political tension around how freely the feature could be used. The debate shifted repeatedly as concerns about budget competed with user value. We landed on a button-triggered approach: users explicitly requested summaries, making the wait time (shown with a skeleton loader) feel intentional rather than frustrating. This solution also enabled intelligent caching—we could store results and avoid expensive redundant API calls until new case activity warranted a fresh summary.

User-triggered AI summary with a deliberate loading state.

Outcome

  • Efficiency Gains: 74% reduction in time required for agents to understand case status, transforming hours of scattered reading into minutes of focused review.
  • Quality Improvements: 45% reduction in customer complaints about receiving incorrect or irrelevant responses from support agents.
  • Technical & Cost Optimization: Reduced unnecessary API calls through intelligent caching and on-demand generation, establishing cost-effective patterns for LLM implementation across the organization.
  • Framework & Standards: Created an evaluation framework for measuring LLM output quality that became the company standard, and pioneered reusable interaction patterns (user-triggered generation, hybrid outputs) adopted by subsequent AI features.

What I learned

  • Technology must serve users, not hype. AI adoption should start with genuine user problems, not the latest capabilities. Connecting technology to real needs prevents building features that impress stakeholders but fail users.
  • Non-deterministic systems demand rigorous evaluation. When outputs are unpredictable, testing with real data and establishing evaluation frameworks isn’t optional—it’s essential to ensure quality and build trust.
  • AI blurs traditional role boundaries. Working with AI technology requires fluid collaboration across disciplines. UX designers must become translators between technical capabilities and human needs, while PMs, developers, and researchers take on hybrid responsibilities.
  • Context is AI’s weakness—and our design opportunity. LLMs struggle with nuanced context in ways humans don’t. Careful prompt engineering and strategic information architecture are critical to getting reliable results.
  • Trust must be designed, not assumed. Users won’t adopt AI features simply because they exist. Transparent interactions, appropriate confidence indicators, and user control are essential to making AI genuinely helpful.

Working on something complex?

I’m always interested in products where systems, workflows and AI create interesting design problems

Contact

LinkedInGithub