AI Coding Evaluator
Helix Labs
Review and rank model-generated code across Python, TypeScript and Go. Flag hallucinations and unsafe patterns.
Longer-term contracts with AI labs and data partners. Activate your account to apply.
16 jobs found
Helix Labs
Review and rank model-generated code across Python, TypeScript and Go. Flag hallucinations and unsafe patterns.
LinguaMind
Translate and quality-check AI training pairs between English and Swahili.
MediGrade AI
Assess clinical accuracy of medical LLM responses. Requires a health science background.
Northbridge
Design evaluation prompts for financial reasoning models and grade outputs.
Sigma Grade
Grade step-by-step model solutions for olympiad-level mathematics.
Sentinel AI
Probe frontier models for unsafe behaviour and document findings.
Quanta Set
Write and validate university-level physics problems with worked solutions.
Vector Ops
Lead a distributed annotation team and own labeling quality metrics.
Fathom AI
Rate AI-generated marketing copy for brand safety, clarity and conversion.
Molecule Bench
Verify reaction mechanisms produced by chemistry models.
NeuralTasks
Entry-level review of mixed AI outputs. Great first role on the platform.
GenomeWorks
Curate and validate biology QA pairs for a research assistant model.
Helix Labs
Write clear documentation and evaluation rubrics for annotation teams.
WaveForm
Transcribe and label multi-accent English audio clips for speech model training.
Scholarly
Compare tutoring model answers against expert rubrics across STEM subjects.
Lexonomy
Annotate contracts and case law for a legal reasoning dataset.