Exploring LLM Evaluations

Investigate AI observability evaluations — hog (deterministic code-based), llmjudge (LLM-prompt-based), and sentiment (user-message sentiment). Find existing evaluations, inspect their configuration, run them against specific generations…

About this skill

Investigate AI observability evaluations — hog (deterministic code-based), llmjudge (LLM-prompt-based), and sentiment (user-message sentiment). Find existing evaluations, inspect their configuration, run them against specific generations, query individual results, and set up scheduled reports on an evaluation.

Maintained by PostHog. The source includes the instructions and any supporting files needed to use this skill.

Inside the instructions

  • 01Tools
  • 02Event schema
  • 03Workflow: investigate why an evaluation is failing
  • 04Step 1 — Find the evaluation
  • 05Step 2 — Break down pass, fail, and N/A
  • 06Step 3 — Read the failing runs

Before you start

  1. Read the instructions and check tool or account requirements.
  2. Install the complete folder when the skill references scripts or other files.
  3. Provide your task context, then review the agent's output.

Source

posthog/ai-plugin / exploring-llm-evaluations

Source reviewed October 2, 2026 · See publisher source