Darum lohnt es sich
• Author detailed, task-based multi-turn conversations and rubrics aligned with project specifications.
• Test and refine conversation drafts against frontier large language models, iterating to meet quality and difficulty requirements.
• Deliver comprehensive evaluation assets including transcripts, target behaviors, binary rubrics, and supporting evidence.
• Ensure strict fidelity to evolving project specs while maintaining high throughput and attention to detail.
• Validate and calibrate outputs with team leads and quality control as guidelines change.
• Work independently and consistently, meeting expected output rates for deliverable completion.
Requirements
• Native-level written English with exceptional clarity, structure, and attention to detail.
• Prior experience in data annotation, RLHF, SFT, evaluation, or prompt engineering for AI systems.
• Working knowledge of frontier LLM behaviors and common model failure patterns.
• Demonstrated ability to interpret and apply highly detailed specifications without supervision.
• Strong critical thinking and analytical skills in writing-heavy or analysis-heavy domains.