Pairwise Preference Rubric (Response A vs Response B)
Evaluates two competitive model completions across strict alignment, instruction following, and conciseness.
Evaluate the two provided assistant completions (Response A and Response B) responding to the user prompt. Evaluation Dimensions: 1. Instruction Adherence: Did the response follow all negative and positive constraints (word count limits, prohibited phrases, required formatting)? 2. Factual Grounding: Are all claims, figures, and technical assertions verifiable without hallucination? 3. Conciseness & Signal-to-Noise: Does the completion deliver direct value without redundant filler or conversational padding? Output Schema: - Preferred Response: [Response A / Response B / Tie] - Justification: Cite explicit quoted evidence from both responses demonstrating why one output is strictly superior. User Prompt & Responses: [INSERT PROMPT AND RESPONSES HERE]