Agent Observability Eval Pipeline

End-to-end Agent Observability pipeline for an instrumented mlapp — classify production traces, root-cause failures, bootstrap evaluators, then (optionally) sample + publish a dataset, generate + run an experiment, and analyze results.…

About this skill

End-to-end Agent Observability pipeline for an instrumented mlapp — classify production traces, root-cause failures, bootstrap evaluators, then (optionally) sample + publish a dataset, generate + run an experiment, and analyze results. Six narrated phases with a standardized banner and a "continue" checkpoint between each. Pure orchestration over the agent-observability sub-skills (agent-observability-session-classify, agent-observability-trace-rca, agent-observability-eval-bootstrap, agent-observability-experiment-bootstrap, agent-observability-experiment-analyzer).

Maintained by Datadog Labs. The source includes the instructions and any supporting files needed to use this skill.

Inside the instructions

  • 01Backend
  • 02Usage
  • 03Inputs
  • 04Precheck
  • 05Phase Template
  • 06Phase 1: Classify ml_app traces

Before you start

  1. Read the instructions and check tool or account requirements.
  2. Install the complete folder when the skill references scripts or other files.
  3. Provide your task context, then review the agent's output.

Source

datadog-labs/agent-skills / agent-observability-eval-pipeline

Source reviewed October 2, 2026 · MIT