Agent Observability Auto Experiment

Run an iterative code-improvement hill-climb against real Datadog LLM-Obs data, locally, with Claude Code as the agent. Establishes a baseline eval, makes one focused change, re-scores with the same harness, keeps the change if it…

About this skill

Run an iterative code-improvement hill-climb against real Datadog LLM-Obs data, locally, with Claude Code as the agent. Establishes a baseline eval, makes one focused change, re-scores with the same harness, keeps the change if it improves the score in the goal's direction (labeling within-noise gains tentative), and repeats.

Maintained by Datadog Labs. The source includes the instructions and any supporting files needed to use this skill.

Inside the instructions

  • 01Security & data handling (read before running)
  • 02Inputs (the experiment config)
  • 03Mandatory intake gate — do this FIRST, before Setup
  • 04Scope — optimize the whole selected surface, not just the prompt
  • 05Domain notes — the product context the code does not carry
  • 06Cost estimate — derived, never asked

Before you start

  1. Read the instructions and check tool or account requirements.
  2. Install the complete folder when the skill references scripts or other files.
  3. Provide your task context, then review the agent's output.

Source

datadog-labs/agent-skills / agent-observability-auto-experiment

Source reviewed October 2, 2026 · MIT