top of page

Agentic IT Operations

B2B enterprise product | 2026

MY CONTRIBUTION

As one of two Product Designers on the project, I helped turn a broad business ambition into a clear product direction by connecting stakeholder goals, operational realities, market shifts, and user needs.
My contribution spanned problem framing, research synthesis, defining human-agent responsibilities and core workflows, information architecture, and prototyping, helping shape a scalable experience that balanced business value, operational efficiency, and user trust.

SCOPE

Product Experience Design, Workflow Design, Detailed Design, Concept Design

SKILS

Product & business thinking, Interviews & observations, Benchmarking & visual research, Research synthesis & concept development, AI-supported exploration, Visual design & prototyping

BACKGROUND

An enterprise IT Operations platform for teams working in complex production environments, designed to help them detect, understand, and resolve operational issues through a unified AI-supported experience.

The platform uses AI agents to connect the dots across fragmented alerts, logs, metrics, services, and incident workflows. It surfaces meaningful issues, explains impact and root cause, and guides teams toward resolution, while keeping agent activity transparent and critical actions under human control.

This case study follows the definition and design of the product’s MVP, from discovery and product framing to detailed UX/UI design, prototyping, and preparation for implementation.

THE PROBLEM

Incident response in production environments is often spread across disconnected dashboards, tickets, logs, alerts, and communication channels. Alerts and signals keep flowing into the system, but they are not automatically connected into a clear incident story.

As a result, human teams are left to build the full picture under pressure: understanding what happened, which signals are related, how severe the issue is, which service is affected, whether there is business or SLA impact, and who needs to take action.

AI agents can reduce much of this manual work, but only if their analysis, conclusions, and actions are visible, understandable, and controlled at the right moments.
Group 340.png

RESEARCH

With this problem in mind, our research focused on understanding the business motivation behind the initiative, how incident response works today across roles, tools, handoffs, and decision points, and where AI agents could meaningfully support the process without reducing transparency or human control.

This helped us identify where users lose time, context, or confidence, and where the product needed to create a clearer balance between automation, trust, and human responsibility.
To build this understanding, we used three research and discovery methods:

UX APPROACH

From doing the work to governing the work

The research shifted our understanding of the problem.
At first, it was easy to frame the opportunity as helping teams resolve incidents faster with AI.
But what became clearer through the research was that incident response was not only fragmented across tools - it was also fragmented across people, handoffs, decisions, and responsibility.

This changed how we approached the product. If agents were going to take on more of the operational work, we first needed to define where humans still create value, and what principles should guide the experience around that shift.

We reframed the human roles for the new operating model, and defined principles to keep the experience clear, evidence-based, and controlled as more responsibility moved from people to agents.
Reframing human roles
We reframed the human roles around the moments where people still create the most value: understanding operational risk, validating technical decisions, approving sensitive actions, and keeping the agents themselves reliable over time.

From there, we defined three key human roles for the new operating model:
image.png

Alex  | Operation / Production Manager

  • Accountable for overall system health and production readiness.

  • Needs to understand operational risk, severity, and business impact.

  • Reviews incidents from a management perspective, not as the hands-on resolver.

  • Approves sensitive production changes when confidence and context are clear.

In the agentic experience, Alex relies on the system to surface early warnings, service health, business impact, SLA risk, and recommended next actions - helping her move from reactive incident assessment to faster, more informed operational decisions.

Design principles
layers.png

Clarity before complexity

Users should first understand what happened, what is affected, how severe it is, and what needs attention. Technical depth should be available, but not compete with the initial operational picture.​

network.png

Trust through evidence

Every recommendation or proposed action should be supported by clear evidence. Users need to understand what the agent checked, what it concluded, and why the action is being suggested before they approve or intervene.

pattern-lock.png

Human control at risk points

The system can automate and recommend, but sensitive operational decisions should remain clear and controllable. Users need visible approval points, manual takeover options, and recovery paths when risk is involved.

THE SOLUTION

From fragmented operations to governed agent work

The solution focused on turning scattered operational signals into a connected workflow: surfacing what needs attention, bringing context and evidence together, clarifying impact and urgency, and creating clear decision points when human input is required.

Agents can take on more of the continuous operational work, while people remain able to understand, validate, approve, and govern the outcome. This supported both sides of the product challenge: helping the business move toward a more scalable and measurable service model, while helping users make faster, clearer, and more confident operational decisions.
Silver-1.png
Frame 2018778311.png

MAIN USER

image.png

Alex
Operation / Production Manager

Monitoring dashboard ​​

The Monitoring Dashboard was designed to help Alex understand the state of production at a glance, before diving into specific incidents or technical details.​

Dashboard - Investigating the Warning-1.png

The dashboard is structured for fast scanning: 

A high-level overview highlights service health, open incidents, SLA risk, and pending approvals, while the service grid organizes production services by severity. Clear status colors, hierarchy, and card sizing help stable services stay quiet and bring at-risk areas forward, so users can quickly focus on what needs attention.

Dashboard - Investigating the Warning - Service panel.png

Service details panel

When a service is selected, the side panel adds focused context without taking the user out of the dashboard.

Instead of showing a generic collection of charts, it highlights the specific KPI that triggered the warning, explains the issue in natural language, and shows the incident raised by the agent. This helps turn raw monitoring data into a clear operational signal - what is happening, why it matters, and what the system is already doing about it.

MAIN USERS

image.png
image.png

Alex
Operation / Production Manager

Daniel 
SRE (Site Reliability Engineer)

Incident details

The Incident Details screen provides a clear, structured view of an incident, bringing together investigation results, context, and next steps in one place so users can understand the situation and decide how to proceed. â€‹

Agents awaiting approval SRE.png

​Instead of forcing users to reconstruct the incident from multiple tools, the screen presents a structured operational narrative: what happened, why it matters, what the agents already investigated, what action was taken or proposed, and what still requires attention.

For Daniel, the screen turns investigation into validation.

He can review the agent’s findings, understand the proposed fix, inspect the reasoning behind it, and decide whether to approve the agent-led action or take over manually when technical judgment is required.

agent activity-1.png

For Alex, the same incident becomes a production decision point. She can review the business impact, risk, expected outcome, and approval context before allowing a sensitive action to move forward. 

business impact.png

When automation needs a human decision

Human intervention is required only when the system cannot resolve independently, or reaches a predefined stop point.

In those moments, the experience creates a clear decision point that explains why input is needed, what outcome to expect, and whether to continue with the agent’s recommendation or take manual control.

Turning resolution into evidence

Once the incident is resolved, the RCA overview closes the loop.

It summarizes the outcome of the incident through resolution time, SLA status, operational cost, and a generated RCA report, turning the resolution process into documented evidence that can support follow-up, learning, and accountability.

root cause analysis.png
By making the full incident lifecycle visible and structured from detection to resolution, the screen turns automation into something users can review, trust, and act on, rather than a black box that simply reports a result.

MAIN USER

image.png

Arjun
FDE  (Forward Deployment Engineer)

Agent reliability management

As AI agents take on more responsibility across incident detection, investigation, communication, and resolution, the product also needs to expose the health of the agent system itself.

For Arjun, the question is no longer only whether production services are healthy, but whether the agents operating those services are reliable, predictable, and governed.

 

The Agentic Dashboard gives Arjun a high-level view of the agent layer across the platform.

It shows whether the system is operational, how agents are performing across domains,

where quality may be degrading, and whether policy gates, cost, or integrations require attention.

​

This creates a starting point for agent supervision: instead of discovering failures only after an agent affects an incident, the user can monitor the reliability of the automation layer itself and identify where deeper investigation is needed.

From domain signal to agent-level investigation

When a domain shows degraded performance or unusual activity, the experience provides a natural path into the Agent Inventory.

 

This view turns the agent layer into a searchable operational list, helping Arjun move from a broad system signal to the specific agent that may require attention.

Agent Inventory.png

From agent list to focused inspection

The Agent Inventory provides a quick way to scan agents, compare their current state, and identify which ones may need attention. Selecting an agent opens a preview panel with the key context needed before moving into the full detail view, helping the user decide whether to keep monitoring, send the agent to sandbox, or continue to deeper management.

Agent Inventory - Open panel.png
Agent Inventory-1.png

The agent control layer

The Agent Control Center gives the FDE a deeper view into a specific agent and the layers that shape its behavior. Instead of treating the agent as hidden automation, the page makes its role, capabilities, dependencies, autonomy level, models, guardrails, activity, and configuration history accessible in one place.

This helps the user understand not only what the agent is doing, but how it is configured, what systems it depends on, how much independence it has, and where its behavior is constrained or needs human oversight. Together, these views make the agent easier to investigate, tune, and govern over time.

OUTCOMES & NEXT STEPS

From MVP definition to production validation

During the MVP phase, the product experience was defined through detailed UX/UI design, prototyping, and preparation for implementation, guided by success criteria established with stakeholders:
efficiency, faster resolution, reliability, trust and control, scalability, and continuous learning.
For users, this meant replacing fragmented incident information with a unified narrative, clear evidence, and visible approval points that enable them to act with confidence under pressure.

The design was presented to stakeholders, product leads, customers, and IT Operations teams, prompting discussion around agent autonomy and human control and supporting the decision to move into implementation.

The next phase is to validate the model in production against these criteria, refine agent autonomy, and expand the product into additional operational workflows.

© 2026 Uria Graiver. All rights reserved.

bottom of page