Core AI Trainer & Quality Lead
Grade AI model responses across conversational realism, logical consistency, and source accuracy.
opportunity
21 AI opportunities
Grade AI model responses across conversational realism, logical consistency, and source accuracy.
Fact-check claims made by foundation models against authoritative primary sources.
Help evaluate and improve AI models using your professional expertise.
Collaborate directly with alignment researchers to audit multi-step reasoning traces and edge-case hallucinations.
Write graduate-level benchmark questions and verify model answers across STEM and legal specialties.
Score candidate model responses for helpfulness, harmlessness, accuracy, and refusal justification.
Audit mathematical reasoning in cutting-edge reasoning models with graduate-level rigor.
Inspect multi-step agent execution traces (browser actions, API calls, spreadsheet manipulations).
Build robust retrieval-augmented generation pipelines and benchmark vector search precision.
Design benchmark suites to measure frontier Gemini model progress on autonomous tool use.
Lead adversarial stress-testing campaigns against upcoming frontier foundation models.
Challenge conversational models with deceptive premises to test whether they resist sycophancy.
Test foundation models for subtle demographic and ideological bias using adversarial test batteries.
Train and calibrate next-generation visual reasoning architectures combining vision and language.
Design autonomous developer agents that write pull requests, execute unit tests, and resolve issues.
Review and calibrate autonomous code generated by fine-tuned models across major web frameworks.
Design supervisory models and automated constitutional rule sets that evaluate agentic workflows.
Review tool-calling outputs generated by enterprise AI assistants to ensure schema validity and security.
Evaluate literary prose, character voices, and narrative flow generated by language models.
Evaluate model translation quality, colloquial idiom comprehension, and culturally sensitive prompts.
Write gold-standard software engineering prompts and rate LLM code generation for runtime safety.
Help evaluate and improve AI models using your professional expertise. Review prompt responses, draft high-quality synthetic training data, and provide domain-informed feedback.
Practical 45-minute evaluation rubric test analyzing 3 foundation model responses.