Back to skills
Agent Launcher

Grade Iterate

Grade Iterate: Improve an agent using a bounded evaluation loop. Review a fixed evaluation set, rubric, baseline results, and failure examples and produce an evaluation record with the next decision.

---
name: agent-launcher-grade-iterate
description: Use for grade iterate when asked to improve an agent using a bounded evaluation loop; produce an evaluation record with the next decision.
license: MIT
metadata:
  author: Thrive
  category: agent-launcher
---

# Grade Iterate

## When to use

Use this skill for grade iterate when you need to improve an agent using a bounded evaluation loop. The expected result is an evaluation record with the next decision.

## Boundaries

Work within the requested task and its stated acceptance criteria. Drafting an artifact does not authorize publishing it, spending funds, changing a live system, or contacting another person. Identify any such action separately before taking it.

## Inputs

Inspect a fixed evaluation set, rubric, baseline results, and failure examples. Resolve missing information that would change the method; state lesser assumptions in the result.

## Method

1. **Diagnose.** Draw the path from user request to tool call to external effect. List the account owner, tool permissions, data stores, rate limits, and actions that cannot be undone.
2. **Decide.** Change one behavior at a time and compare it with the baseline.
3. **Produce.** Build an evaluation record with the next decision from the inspected material; keep assumptions distinguishable from observed facts.

## Decision rules

- Gate launch on a bounded task, a measurable success condition, an explicit stop rule, and a fallback for missing context or tool failure. Keep the first run supervised.
- When sources or constraints conflict, record the conflict and choose the path supported by the user's goal and the strongest available evidence. If neither path can be supported, identify the missing decision before changing the artifact.

## Domain rules

- Set tool permissions, spending limits, and a human-owned stop condition before launch.
- Evaluate a complete task and a failed run before allowing recurring execution.

## Verification

Stop when the target passes or the iteration limit is reached. Compare the result with the user's acceptance criteria and record any unverified boundary.

Record an example task, the agent trace, expected result, actual result, cost, and the human decision required before recurring execution.

## Approval and stop conditions

- Before recurring or external actions, identify the owner, tool permissions, spending limit, approval gate, and disable path.
- Stop when the agent cannot determine whether an action succeeded, reaches its retry or budget limit, or would exceed an agreed boundary.
- Pilot with human review before scheduling unattended execution. Report unresolved failures and who can disable the workflow.

## Stop conditions

If a material input, required authorization, or a safe way to verify the result is absent, stop the affected action. Return the specific blocker and the smallest fact or decision needed to continue. Do not report an unrun check as passed.

## Output

Provide an evaluation record with the next decision. Include the decisive evidence and actual verification result. Name any artifact location and unresolved issue that affects its use.

<!--
MIT License

Copyright (c) 2026 Thrive

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
-->