The short version
- Read the rubric → Check evidence → Explain the rating.
- Keep source evidence and review the result before using it.
An AI evaluator reviews model outputs against a defined task and records a reason for the judgment. The work can involve checking facts, comparing answers, reviewing safety issues, or testing whether an assistant followed instructions. The rubric and employer determine the exact assignment.
A useful evaluation explains why an answer passes or fails. "This one sounds better" does not give another reviewer enough information to repeat your decision.
Follow the assignment from input to decision
Stage | What the evaluator does | What another reviewer should be able to inspect |
|---|---|---|
Understand the task | Read the question, source material, and instructions | The intended answer and task boundaries |
Apply the rubric | Check each required criterion | The criterion supporting each rating |
Verify evidence | Compare claims with an authorized reference | The source location and any uncertainty |
Record the decision | Explain a preference, tie, or failure | A concise reason linked to observable content |
Escalate an exception | Flag an unclear rule or sensitive case | The unresolved issue and required review |

Keep factual accuracy, instruction following, completeness, and writing style separate when the rubric allows it. A readable response can contain a serious factual error. A correct response can still omit a required field.
Try a fictional evaluation exercise
Imagine a source note says: "The workshop starts at 10:00 on Tuesday. Registration closes on Monday. The room has not been assigned."
Response A says: "The workshop is Tuesday at 10:00 in Room 4. Register by Monday."
Response B says: "The workshop starts Tuesday at 10:00. Registration closes Monday. The room is not specified."
Criterion | Response A | Response B |
|---|---|---|
Preserves the stated date and time | Yes | Yes |
Includes the registration deadline | Yes | Yes |
Avoids inventing a room | No; Room 4 is unsupported | Yes; marks the missing information |
Overall decision for this exercise | Reject for an unsupported factual detail | Prefer because it preserves the source boundary |
The example is fictional and uses a simple rubric. A real project may define different labels and severity levels. Apply those rules rather than copying this decision scale into an unrelated assignment.
Write reasons another person can verify
Quote the relevant part of the response when the project's rules permit it. Identify the criterion and describe the issue. For the exercise above, "invented Room 4, which the source does not assign" is more useful than "hallucination."
If the reference material is unclear, say so. Do not fill the gap with a web search unless the task authorizes external research. Record whether you found a factual conflict, a missing reference, or a rule that needs clarification.
The pairwise evaluation prompt provides a practice framework. Confirm its instructions against the rubric you are asked to use. The model evaluation workflow can help organize a review log.
Resolve disagreements without hiding them
Two reviewers may agree on an error and disagree on its severity. Keep both observations. Ask which rubric rule distinguishes a minor issue from a failure, then record the clarification for the next batch.
A calibration exercise should include ordinary cases and difficult ones. Review whether the instructions are precise enough to support consistent judgments. A tie or escalation can be a valid outcome when the evidence does not justify a preference.

Understand the project's conditions before accepting work
Read the data handling rules, permitted tools, expected availability, and payment terms. Ask how review disputes and rejected work are handled. If the task involves sensitive material, understand the support and escalation options before beginning.
Outlier describes rating answers and creating rubrics among its public task examples. Its opportunity requirements vary. These examples establish that the tasks exist; they do not establish your eligibility, a current opening, or expected earnings.
Show your judgment in a portfolio
Create a small practice set using public or invented material. Include the task, rubric, reference, responses, ratings, and explanations. Mark which examples you wrote yourself. Remove personal data and material you lack permission to publish.
On a resume, describe the scope honestly: "Compared six practice responses against a source accuracy rubric and documented disagreements" is a specific project claim. Do not say you improved a commercial model unless you have evidence for that contribution.
Read the guide to getting a job in AI, then prepare a relevant version in the resume builder. The AI evaluation hub contains other practice workflows.
Sources and review notes
Thrive Editorial reviewed these sources on September 28, 2026. The workshop exercise and ratings are original teaching material. This guide describes evaluation work and does not certify a reader's proficiency or promise employment.
Put it into practice
Your next step
Have a question or a correction?
Contact Thrive


