Model Output Error Taxonomy & Concordance Audit
Categorizes model failure modes across a dataset into structured error buckets and computes agreement rates.
You are a Principal AI Evaluation Lead. Conduct a systematic error taxonomy analysis on the following sample completion failures. Classification Rules: 1. Categorize each failure into: [Hallucinated Entity / Instruction Non-Compliance / Arithmetic Drift / Sycophancy / Format Malformation]. 2. Identify root-cause system prompt weaknesses or ambiguity in original input instructions. 3. Compute an error distribution percentage across the sample batch. 4. Propose 3 prompt-engineering guardrails or negative constraints to eliminate at least 80% of observed failure modes. Sample Dataset: [INSERT DATASET HERE]