Apu P.
AI Evaluation Specialist | Data Annotation | Human-in-the-Loop
Is your AI generating responses that sound correct but still contain mistakes? I can help you identify them through careful human evaluation. I provide AI/LLM evaluation, data annotation, classification, and quality review based on your specific guidelines and evaluation criteria. I can help identify: -Factual errors and incorrect information -Irrelevant or incomplete responses -Incorrect classifications and false positives -Unsupported reasoning or conclusions -Tone and context issues -Edge cases and guideline violations I don't blindly accept AI outputs. I carefully compare the response with the original prompt, source information, context, and your rubric before making a judgment. When a case is genuinely unclear, I flag it rather than guessing. Whether you're building an LLM, training an AI model, creating a dataset, or performing AI quality assurance, I can provide reliable human evaluation to help improve your system. โณ Availability ๐ 40+ hr/week ๐ Open to long-term or project-based roles If you're ready to remove overwhelm, stay organized, and get more done โ letโs connect.