Telemetry, as a Conversation

Moving telemetry from a specialist tool into everyday conversation - so that the users can ask, verify, and share data without leaving their workflow

Role

Product Designer

Timeline

Oct 2025 - Feb 2026

Scope

0 → 1

Domain

Conversational AI, Data Accessibility, Enterprise Tooling

Due to NDA, certain details, data, and visuals have been simplified or abstracted to respect confidentiality. The case study focuses on design process and decision-making.

# Overview

At a Glance

The organization's analytics platform wasn't broken, it was just dense; internal shorthand, hover-only tooltips, tables you had to scroll to read. Daily users, mostly PMs and engineers, had learned to route around it; everyone else was left asking them instead. For designers specifically, that meant getting data secondhand - staying out of the interpretation loop entirely.


This project embedded a conversational AI assistant into Microsoft Teams - one that understands plain-language questions and turns them into personalized, decision-ready dashboards.


Led research, design, and prototyping end-to-end - partnering with a Product Manager and three engineers on scoping and implementation, with periodic feedback from a Senior Product Designer.

# Context

A System People Adapted To, Not One That Worked

The organization had a capable analytics platform - dashboards, filters, drill-downs, everything engineers and PMs needed day to day. What it didn't have was a design that assumed a first-time user. A quick heuristic pass made the pattern concrete: this wasn't a designer problem, it was a "you had to already know" problem, and it applied to anyone who didn't use the tool constantly.

Learnt, Not Labeled

Filter names like "Segment X" or "Type Y" only made sense if you already knew what they stood for - shorthand daily users absorbed through repetition, with no way in for anyone else.

Filter labels relied on shorthand only daily users understood, absorbed through repetition, with no way in for anyone else.

Hidden by Default

Metric definitions existed, but only behind hover states most people never found. One filter panel needed its own inline note just to explain its own default.


Metric definitions existed, but only behind hover states most people never found. One filter panel needed its own inline note just to explain its own default.

Costly to Relearn

Every visit meant re-figuring out which client, which segment, which of two near-identical filters to pick. Daily users absorbed that cost into habit; everyone else just asked someone instead.

THE REALITY

  • Has a data question, but can't dig in alone


  • Pings a PM or engineer, or catches them in a meeting


  • Waits on their availability - sometimes days


  • Gets a secondhand answer, occasionally missing context or slightly off


  • Any follow-up question means starting the cycle over

WHAT’S NEEDED

  • Types a question directly in Microsoft Teams


  • AI parses intent and generates a validated query


  • Dashboard preview appears in under 2 minutes, with metrics defined inline


  • One-click share in the same thread


  • Refines or forks the dashboard directly - no one else needed

  • Types a question directly in Teams


  • AI parses intent and generates a validated query


  • Dashboard preview appears in under 2 minutes, with metrics defined inline


  • One-click share in the same thread


  • Refines or forks the dashboard directly - no one else needed




# The Approach

Which fix actually worked? ...the one that didn't ask anyone to open a new tool.

Two other directions got tested before this one: redesigning the existing dashboard, and building a standalone query-builder app. Both solved a piece of the problem, but neither solved where the friction actually lived, people weren't in the telemetry tool often enough to make either fix matter. The answer was to stop trying to fix the tool and meet people where they already worked instead. Here's the research that surfaced that gap, and what got explored before landing here.

# Secondary Research

Understanding the Gap Between Data and Decisions

Before designing solutions, the priority was understanding why a technically functional tool was going unused - and whether the barrier was capability, confidence, or context.

14

Contextual Interviews

10 UX designers, 2 product managers, 2 engineers - observed in their environment to surface real workflows and workarounds.

4

Usability Tests

Tested the existing platform across roles to isolate friction: onboarding, query construction, interpreting results, sharing.

20

Cross-Functional Survey

Surveyed product, design, and analytics teams on data access and trust: 18 of 20 default to asking a PM or engineer over using the platform directly.

3

Stakeholder Workshops

Sessions with analytics and infrastructure teams to map technical constraints, governance requirements, and data pipeline realities.

Key Insights

1

Fluency was earned through repetition, not granted by role

PMs and engineers weren't given a better tool - they'd logged enough hours in it to route around unclear labels. That fluency never transferred to occasional users.

2

1

The barrier was cognitive load, not permission

Fluency was earned through repetition, not granted by role

Everyone had access. Internal shorthand, hover-only definitions, and dense tables made it unusable without months of accumulated context.

PMs and engineers weren't given a better tool - they'd logged enough hours in it to route around unclear labels. That fluency never transferred to occasional users.

2

The barrier was cognitive load, not permission

Everyone had access. Internal shorthand, hover-only definitions, and dense tables made it unusable without months of accumulated context.

3

Designers think in questions, not queries

They ask 'what are people searching for?' - the tool required navigating filter taxonomies built for a different mental model.

4

Low trust in the data

Inconsistent definitions made people hesitant to cite data confidently in reviews or share it with stakeholders.

5

Sharing lost context

A shared dashboard was just a link - the insight behind it rarely traveled with it.

3

Designers think in questions, not queries

They ask 'what are people searching for?' - the tool required navigating filter taxonomies built for a different mental model.

6

Where the work already happened

Most participants spent their working hours in Microsoft Teams, not the telemetry platform. Any real fix needed to meet them there.

4

Low trust in the data

Inconsistent definitions made people hesitant to cite data confidently in reviews or share it with stakeholders.

5

Sharing lost context

A shared dashboard was just a link - the insight behind it rarely traveled with it.

6

Where the work already happened

Most participants spent their working hours in Microsoft Teams, not the telemetry platform. Any real fix needed to meet them there.

Design Hypothesis

If telemetry access is embedded into the collaboration layer as a conversational interface - with transparent metrics and one-click sharing - people will self-serve data questions without analyst mediation, increasing both the frequency and quality of data-informed decisions.

# Exploration

Three Directions, One Real Answer

Before landing on an embedded conversational assistant, two other directions were explored and set aside - each for a specific, testable reason.

Option A - Redesign the existing dashboard

Fixed the interface, not the underlying mental-model mismatch interviews had already surfaced.


Option B - A standalone query-builder app

Fixed the query-syntax problem, but reintroduced the exact context-switching cost which research had identified as the real barrier.

Embedded conversational assistant (chosen)

Met people in the workflow they already used, removed the query-syntax barrier, and made sharing a natural extension of how teams already talked about work.

# Solution

A Telemetry Assistant That Lives Where Teams Collaborate

Rather than improving the existing dashboard tool, the solution embedded telemetry into Microsoft Teams as a conversational AI assistant - transforming data access from a destination to a capability. Three screens show the critical path.

Due to NDA constraints, the actual production screens cannot be shared. These were recreated for the purpose of this case study using AI-assisted design, reflecting the same design decisions and workflows.

Transparency

Before fetching any data, the assistant shows its interpretation - scope, date range, metrics, and sort order - so users can verify intent and correct errors before execution. This single decision was the core trust mechanism: it eliminates a full class of errors without requiring technical knowledge.

Transparency

Before fetching any data, the assistant shows its interpretation - scope, date range, metrics, and sort order - so users can verify intent and correct errors before execution. This single decision was the core trust mechanism: it eliminates a full class of errors without requiring technical knowledge.

Results

A rich card delivers data, charts, and AI-generated observations together. The observations shift users from looking at data to understanding implications - flagging that 25% of searches return no results is a usability signal designers can act on immediately, without interpreting raw numbers themselves.

Edge Case

When a query is genuinely ambiguous, the assistant surfaces one focused clarifying question with selectable options - never an open prompt, never a failure state. The constraint of one question only was deliberate: testing showed multi-question clarification flows caused query abandonment.

Edge Case

When a query is genuinely ambiguous, the assistant surfaces one focused clarifying question with selectable options - never an open prompt, never a failure state. The constraint of one question only was deliberate: testing showed multi-question clarification flows caused query abandonment.

# Key Decisions

Why This, Not That

Every choice was evaluated against one principle: reduce cognitive load while increasing trust. Each decision was a trade-off, not a default.

Why embed in Teams rather than improve the existing tool?


That's where people already worked. No new tool to adopt, no context-switching.

Why progressive disclosure for the SQL preview?


Technical users want to see the query; everyone else finds it intimidating. Hidden by default, one click away.

Why show query interpretation before execution?


Cheaper to prevent errors than fix them - people verify intent before data is fetched.

Why suggested queries on the home screen?


A blank state killed early adoption. Example queries lowered the barrier - tooltips got skipped entirely.

Why AI-generated observations alongside raw data?


Raw numbers still need interpreting. A flagged signal like "25% no results" makes the insight usable instantly.

Why collaborative sharing within Teams?


Data discussions happen in threads, not links. Sharing in-place kept context intact.

# Edge cases

AI + Data Demands Credibility

Conversational AI in a data context carries unique trust risks. Two failure states required deliberate design rather than fallback behavior.

Ambiguous Intent

Failure State

Rather than guessing or erroring out, the assistant asks one focused clarifying question with selectable options - no need to re-enter the query.


For eg. Choosing one structured question over an open prompt: Open prompts caused abandonment in testing; selectable options kept people in the flow.

Low Data Confidence

Trust Risk

When results are sparse or the scope too broad, the assistant flags it outright rather than presenting thin data as complete. Metric definitions - owner, source, last updated - sit inline, at the point of use.

For eg. Choosing inline provenance over separate docs: No one has to leave the conversation to understand what a metric means.


Trust mechanisms to be built into the system

Query interpretation preview

SQL preview (expandable)

Metric provenance inline

Data freshness timestamp

One clarifying question max

Governance by default (SSO, PII masking)

# Outcomes

Measuring What Matters

These metrics were defined during research, before launch. A team demo of the core flow - interpretation preview, results card, clarifying question - was well received, which gave confidence to move forward. Development is currently paused, with resumption expected.

Confidence in data

Are people citing data more often in reviews and discussions - without waiting on someone else to pull it?

Trust in AI outputs

Do people trust results enough to share without double-checking first? Does seeing the query change that?

Request volume to PMs/engineers

Is direct-ask volume dropping? Baseline set during research; to be re-measured once development resumes and the assistant ships.

Data discussion patterns

Are conversations shifting from mediated to self-served - and does context still travel with the data when shared?

# What's next

After Launch

Next up: role-based suggested queries, personalizing the home screen for different data needs. Longer term - anomaly detection alerts, and fine-tuning the model on internal telemetry vocabulary to sharpen intent parsing.

# Reflections

From Designing Tools to Designing Capabilities

1

The most important design decision was the reframe.

Improving the tool would've made a hard thing slightly less hard. Moving it into Teams made it invisible - no new system to learn. That only became visible once the actual workflow was understood, not assumed.

2

Trust in AI systems is a design surface, not a feature.

Query transparency, metric provenance, clarifying questions - each was a trust decision, not a UX detail. Showing scope before execution looked minor; it decided whether people trusted the results at all.

3

Compliance and security should have been in research, not review.

The trust framework was built from a designer's view alone. Involving compliance and security as research participants, not reviewers, would have caught governance constraints early - a design advantage instead of late correction.

The most important design decision was the reframe.

Improving the tool would've made a hard thing slightly less hard. Moving it into Teams made it invisible - no new system to learn. That only became visible once the actual workflow was understood, not assumed.

Trust in AI systems is a design surface, not a feature.

Query transparency, metric provenance, clarifying questions - each was a trust decision, not a UX detail. Showing scope before execution looked minor; it decided whether people trusted the results at all.

Compliance and security should have been in research, not review.

The trust framework was built from a designer's view alone. Involving compliance and security as research participants, not reviewers, would have caught governance constraints early - a design advantage instead of late correction.

Create a free website with Framer, the website builder loved by startups, designers and agencies.