Eval Engineering

Inspect an agent repository and optional traces, interview the user, write reviewed Task Specs, build and audit Harbor tasks, and bootstrap reusable project World Knowledge Skills. Use for agent evals, benchmark design, Task generation…

About this skill

Inspect an agent repository and optional traces, interview the user, write reviewed Task Specs, build and audit Harbor tasks, and bootstrap reusable project World Knowledge Skills. Use for agent evals, benchmark design, Task generation, controlled Environments, synthetic data, Verifiers, Harbor runs, calibration, or continuous benchmark maintenance.

Maintained by LangChain. The source includes the instructions and any supporting files needed to use this skill.

Inside the instructions

  • 01Flow
  • 02Terms
  • 03Reference routing
  • 041. Inspect inputs and existing World knowledge
  • 052. Propose and select a Task
  • 063. Write and review the Task Spec and World Skill

Before you start

  1. Read the instructions and check tool or account requirements.
  2. Install the complete folder when the skill references scripts or other files.
  3. Provide your task context, then review the agent's output.

Source

langchain-ai/langchain-skills / eval-engineering

Source reviewed October 2, 2026 · MIT