Find your next AI

opportunity

Browse jobs

21 AI opportunities

Core AI Trainer & Quality Lead

DataAnnotationDataAnnotation•Worldwide Remote
$30–$60/hrFlexibleNo experience

Grade AI model responses across conversational realism, logical consistency, and source accuracy.

Today

General Knowledge & Factuality Verifier

DataAnnotationDataAnnotation•Worldwide Remote
$25–$45/hrFlexibleNo experience

Fact-check claims made by foundation models against authoritative primary sources.

Today

AI Training Specialist

MercorMercor•Remote
$40–$120/hrFlexibleExpert

Help evaluate and improve AI models using your professional expertise.

Today

AI Reasoning & Alignment Researcher

OpenAIOpenAI•San Francisco, CA
$80–$140/hrFlexibleExpert

Collaborate directly with alignment researchers to audit multi-step reasoning traces and edge-case hallucinations.

Today

Domain Expert Contributor (STEM & Law)

MercorMercor•Global Remote
$50–$120/hrFlexibleExpert

Write graduate-level benchmark questions and verify model answers across STEM and legal specialties.

Today

RLHF Prompt & Safety Evaluator

AnthropicAnthropic•Remote
$55–$95/hrFlexibleMid

Score candidate model responses for helpfulness, harmlessness, accuracy, and refusal justification.

Today

Advanced Mathematics & Proof Evaluator

OutlierOutlier•Remote
$65–$120/hrFlexibleExpert

Audit mathematical reasoning in cutting-edge reasoning models with graduate-level rigor.

Today

Autonomous Agent Behavior Evaluator

Scale AI•Worldwide Remote
$40–$75/hrPart-timeEntry

Inspect multi-step agent execution traces (browser actions, API calls, spreadsheet manipulations).

Today

Open-Source Agent Pipeline Specialist

Micro1Micro1•Remote Global
$55–$95/hrPart-timeMid

Build robust retrieval-augmented generation pipelines and benchmark vector search precision.

Today

Benchmark & Evaluation Specialist

Google DeepMindGoogle DeepMind•London / Mountain View
$210k–$340k/yrFull-timeSenior

Design benchmark suites to measure frontier Gemini model progress on autonomous tool use.

Today

Frontier Model Red Teamer & Bias Evaluator

OpenAIOpenAI•San Francisco, CA / Hybrid
$220k–$350k/yrFull-timeExpert

Lead adversarial stress-testing campaigns against upcoming frontier foundation models.

Today

Truth & Reasoning Adversarial Tester

AnthropicAnthropic•Remote
$60–$105/hrPart-timeMid

Challenge conversational models with deceptive premises to test whether they resist sycophancy.

Today

AI Model Alignment & Bias Auditor

EthosEthos•Remote
$50–$90/hrFlexibleMid

Test foundation models for subtle demographic and ideological bias using adversarial test batteries.

1d ago

Multimodal AI Research Engineer

Google DeepMindGoogle DeepMind•London / Mountain View
$240k–$380k/yrFull-timeExpert

Train and calibrate next-generation visual reasoning architectures combining vision and language.

1d ago

Fullstack Agentic Pipeline Engineer

MercorMercor•Global Remote
$80–$150/hrFull-timeSenior

Design autonomous developer agents that write pull requests, execute unit tests, and resolve issues.

1d ago

AI Technical Vetting Contributor

Micro1Micro1•Worldwide Remote
$45–$85/hrPart-timeMid

Review and calibrate autonomous code generated by fine-tuned models across major web frameworks.

Yesterday

AI Alignment & Constitutional Evaluator

AnthropicAnthropic•San Francisco, CA / Remote
$190k–$320k/yrFull-timeSenior

Design supervisory models and automated constitutional rule sets that evaluate agentic workflows.

Yesterday

Enterprise RAG & Tool-Use Evaluator

EthosEthos•Remote
$65–$110/hrFlexibleMid

Review tool-calling outputs generated by enterprise AI assistants to ensure schema validity and security.

2d ago

Creative & Fiction Tone Alignment Specialist

AlignerrAlignerr•Remote
$35–$60/hrFlexibleEntry

Evaluate literary prose, character voices, and narrative flow generated by language models.

2d ago

Multilingual LLM Reasoning Reviewer

OutlierOutlier•Remote
$35–$65/hrFlexibleEntry

Evaluate model translation quality, colloquial idiom comprehension, and culturally sensitive prompts.

3d ago

Expert Code & Technical Reviewer

AlignerrAlignerr•Remote
$60–$115/hrFlexibleSenior

Write gold-standard software engineering prompts and rate LLM code generation for runtime safety.

4d ago