Darum lohnt es sich
Responsibilities
• Develop complex, adversarial multi-turn conversations and task-based scenarios aligned with project specifications.
• Author clear, precise evaluation rubrics to assess model responses against defined behavioral targets.
• Iteratively test conversations and tasks against frontier LLMs, escalating difficulty and nuance.
• Deliver comprehensive task packages, including transcripts, target behaviors, and supporting rationale.
• Validate LLM outputs, documenting model strengths and failure modes relative to specifications.
• Maintain calibration with team leads and quality control contacts as project requirements evolve.
Requirements
• Exceptional written English skills with clarity, precision, and strong structural organization.
• Prior experience in AI human data environments such as RLHF, SFT, evaluations, or prompt engineering.
• Deep familiarity with large language models and ability to identify common failure patterns.
• Demonstrated ability to work autonomously, interpreting and executing complex specifications.
• Proven critical thinking and meticulous attention to detail.
• Experience designing evaluation items or rubrics is advantageous.