Darum lohnt es sich
Role Overview
Mercor is partnering with a leading AI research organization to engage experienced data scientists for a project focused on evaluating how well AI systems perform real-world data science work.
Rather than producing deliverables yourself, you will define what excellent work looks like: designing task-specific grading criteria and scoring completed work samples with rigorous, well-reasoned written justifications.
2.
Key Responsibilities
• Design precise, task-specific grading criteria for real-world data science deliverables (analyses, models, dashboards, experiment readouts, and written recommendations)
• Score AI-generated and human work samples against those criteria, with detailed written justifications for every score
• Apply consistent, evidence-based judgment so that scores are reproducible and defensible
• Incorporate structured feedback from senior reviewers and iterate quickly on your work
3.
Ideal Qualifications
• 5+ years of professional data science experience in industry
• Background in business operations, product, or growth data science at top-tier technology companies
• Deep fluency in experiment design and A/B testing, metric definition, SQL/Python analysis, and communicating findings to executive stakeholders
• Exceptionally strong written communication
• Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers
• Prior experience with AI training, evaluation, or human-data projects is a strong plus
4.
Application Process
• Qualified applicants may be asked to complete a brief technical assessment or submit additional information