Agent Observability Eval Bootstrap

Bootstrap evaluators from production traces — by default propose online LLM-judge evaluators and, after you confirm, create them in Datadog as disabled drafts (never auto-enabled); on request emit Python SDK code or a framework-agnostic…

About this skill

Bootstrap evaluators from production traces — by default propose online LLM-judge evaluators and, after you confirm, create them in Datadog as disabled drafts (never auto-enabled); on request emit Python SDK code or a framework-agnostic JSON spec instead.

Maintained by Datadog Labs. The source includes the instructions and any supporting files needed to use this skill.

Inside the instructions

  • 01Backend
  • 02Usage
  • 03Inputs
  • 04Available Tools
  • 05Key get_llmobs_span_content Patterns
  • 06How to Use search_llmobs_spans

Before you start

  1. Read the instructions and check tool or account requirements.
  2. Install the complete folder when the skill references scripts or other files.
  3. Provide your task context, then review the agent's output.

Source

datadog-labs/agent-skills / agent-observability-eval-bootstrap

Source reviewed October 2, 2026 · MIT