Data Trainer with hands-on experience in LLM and Generative AI evaluation, RLHF, multimodal content assessment, and AI-generated translation evaluation. Skilled in analyzing and ranking model outputs for quality, accuracy, naturalness, safety, and policy compliance, including English-to-Portuguese content. Background in language data annotation and data quality, with international professional experience and fluency in English and Portuguese.
• Evaluate and annotate LLM and generative AI outputs across text, image, and video, assessing accuracy, relevance, quality, safety, and alignment with detailed evaluation guidelines.
• Assess AI-generated English-to-Portuguese video translations, comparing original and translated versions for linguistic accuracy, fluency, voice naturalness, word choice, and overall human-like delivery.
• Perform high-judgment evaluation and comparative ranking of AI-generated responses, identifying hallucinations, quality issues, policy violations, and unsafe model behavior.
• Evaluate sensitive and non-sensitive AI-generated content for safety, policy compliance, and legal requirements, applying complex guidelines consistently across diverse tasks.
• Contribute to RLHF (Reinforcement Learning from Human Feedback) workflows through structured human evaluation, classification, annotation, and ranking to support model quality and alignment.