Agent Observability Build Eval From Annotations

Fit a Datadog LLM-Obs evaluator to human labels. Takes an annotation queue, works out where in the trace the labelled property actually lives, drafts an LLM-judge that predicts the human label, scores that judge against the…

About this skill

Fit a Datadog LLM-Obs evaluator to human labels. Takes an annotation queue, works out where in the trace the labelled property actually lives, drafts an LLM-judge that predicts the human label, scores that judge against the already-labelled rows with a metric agreed with the user, then hill-climbs it — inspect the errors, make one focused change, re-score, keep it only if it beats the best — for a bounded number of iterations, and finally publishes the winner to Datadog as a DISABLED evaluator (not a Datadog draft — a real evaluator with enabled: false).

Maintained by Datadog Labs. The source includes the instructions and any supporting files needed to use this skill.

Inside the instructions

  • 01Security & data handling (read before running)
  • 02Inputs
  • 03Intake gate — before anything else
  • 04Datadog backend — MCP or pup
  • 05State — .build_eval_from_annotations/
  • 06Phase 1 — Read the queue and the labels

Before you start

  1. Read the instructions and check tool or account requirements.
  2. Install the complete folder when the skill references scripts or other files.
  3. Provide your task context, then review the agent's output.

Source

datadog-labs/agent-skills / agent-observability-build-eval-from-annotations

Source reviewed October 2, 2026 · MIT