THE LINUX FOUNDATION PROJECTS

WEBINAR | Observing Agentic AI: A New Category of Development and Operational Problem

Your AI agents are reasoning in the dark.

And most teams can’t see why.

When an agent hallucinates, mis-plans, or burns through a multi-pass reasoning trace, “it was the model” isn’t a post-mortem — it’s a blind spot.

This session is for engineering leaders who need decision transparency, not just uptime.

ABOUT THE WEBINAR

As enterprises shift to agentic AI systems, the development and operational surface area is fundamentally changing. Unlike deterministic software, AI agents reason, plan, and iteratively execute multi-step workflows to arrive at a complete response. This evolution introduces a new category of operational problems: how do you monitor a system that can “hallucinate,” troubleshoot a reasoning trace, or audit a non-deterministic decision?

This challenge extends to development practices as well, where software engineers struggle to build agents without clear observability into their accuracy, performance, and cost. Traditional observability frameworks focused on “RED” metrics (Rate, Errors, Duration) are no longer sufficient. Today’s engineering leaders must manage agentic health, tracking token consumption, latency across multi-pass reasoning, and the accuracy of automated actions.

In this session, we will explore how the OpenSearch Observability Stack, a unified, vendor-neutral, and Apache 2.0-licensed platform, solves these complex challenges. We will discuss how to move from a “data swamp” to actionable intelligence by correlating logs, traces, and Prometheus metrics to ensure agentic stability at scale.

Join Dotan Horovits and Megha Goyal from AWS as we bridge the gap between raw telemetry and strategic insight, using OpenSearch as the neutral, scalable engine for observing the next generation of AI.

KEY TAKEAWAYS

The Observability Gap

Understanding why agentic AI requires a shift from infrastructure monitoring to decision transparency, including result explainability and confidence signaling.

Unified Telemetry at Scale

How to leverage OpenSearch to correlate distributed traces with logs and metrics, reducing Mean Time to Resolution (MTTR) for complex AI pipelines.

Protecting the “Brain”

Managing the sensitive data generated by agent interactions while maintaining data sovereignty and avoiding proprietary vendor lock-in.

Operational ROI

Strategies for balancing model complexity with cost-effective, high-performing storage and indexing for petabyte-scale telemetry.

SPEAKER

Bridging the gap between raw telemetry and strategic insight for the next generation of AI.

Dotan Horovits

Senior Developer Advocate, AWS

Megha Goyal

Senior Software Engineer, AWS

“How do you monitor a system that can hallucinate, troubleshoot a reasoning trace, or audit a non-deterministic decision?”