Darum lohnt es sich
Key Responsibilities
• Execute structured annotation tasks: classification, ranking, span labeling, and named entity recognition (NER)
• Perform LLM evaluation and prompt evaluation (helpfulness, factuality, groundedness, instruction-following)
• Contribute to RLHF workflows via preference judgments, rationale writing, and comparative response evaluation
• Run QA evaluation through self-checks, peer reviews, and audit responses
• Document edge cases, track error patterns, and propose guideline clarifications to reduce variance
Projects and Task Types
• Multilingual NER and text labeling
• Content safety labeling aligned to policy taxonomies
• Computer vision annotation (bounding boxes, segmentation, image-text alignment) as needed
• Calibration sets and inter-review agreement improvement activities
Required Qualifications
• Experience with guideline-based labeling, evaluation frameworks, QA, data operations, or similar rubric-driven work
• Ability to interpret complex guidelines and apply consistent decisions at scale
• Comfort with web-based annotation tools, task queues, and documentation of edge cases
• Strong written English for rationale-based RLHF comparisons and clear communication
Benefits
• Competitive hourly rate: $30–$50/hr
• Fully remote, full-time structured workflows with calibration and quality checkpoints
• Work on real-world LLM evaluation and multimodal annotation programs Remote Data Annotation Jobs Madrid (Full-Time, Remote)
Join production-grade AI/ML training workflows by labeling and evaluating data used to improve model performance across NLP, computer vision, and content safety systems.
You will follow strict rubrics and maintain strong annotation guidelines compliance to protect training data quality for LLM training pipelines.