AI Evaluation Analyst

Curvia AI
Remote, Brazil
Publicada: 23/09/2026
Via: whatjobs

Descrição da Vaga

Company Description:

Curvia AI is a London, UK-based technology and AI solutions company that transforms complexity into opportunity through enterprise solutions, AI-powered automation, and data-driven insights. We build practical, scalable, and human-centered technology grounded in integrity, security, and responsible innovation.


Mandatory Req

  • 5+ yrs of relevant Experience
  • Comfortable reading Python, SQL, shell scripts, structured data, and execution logs to understand task setup and grading behavior.


Offer Details:

  • Pay : US $100 per day (With 8 approved hours per day)
  • Mode of work: Fully Remote
  • Number of positions: 25
  • Mode of Employment: Contract (no medical/paid leave)
  • Commitment Required : Full-time. 40 hours per week with 8 hours PST overlap per day
  • Experience : 5+ years
  • Start Date: Immediate


What we're looking for:

  • Technical fluency: Comfortable reading Python, SQL, shell scripts, structured data, and execution logs to understand task setup and grading behavior.
  • Analytical judgment: Able to independently check calculations, reconcile conflicting evidence, and assess the correctness and completeness of professional deliverables.
  • Clear communication: Strong written English, attention to detail, and experience providing specific, reproducible feedback.
  • Relevant experience preferred: AI evaluation, technical QA, data analysis, or benchmark development; familiarity with Harbor task setup. 


Responsibilities: 

  • Validate task quality: Check that instructions, source materials, reference solutions, and evaluation criteria are consistent, with no hidden requirements or missing information.
  • Review agent performance: Inspect execution traces, tool calls, and generated deliverables to determine whether successes and failures are justified.
  • Audit grading logic: Identify brittle checks, incorrect expected answers, unsupported rubric criteria, and cases where valid alternative solutions are unfairly penalized.
  • Investigate discrepancies: Distinguish genuine model limitations from task defects, grader errors, and environment or tool failures. Assess automated QC findings independently rather than accepting them at face value.
  • Document decisions: Provide concise, evidence-backed findings and actionable feedback, flag uncertainty, and verify that revisions resolve identified issues.



Interessado nesta vaga?

Candidate-se agora pelo site oficial da empresa

Informações da Vaga

  • Local Remote, Brazil
  • Publicada 3 horas atrás

Sobre a Empresa

Curvia AI

Vagas Semelhantes