You are viewing a preview of this job. Log in or register to view more details about this job.

Freelance Academic Research Expert AI Evaluator

Uber AI Solutions is Uber's marketplace connecting skilled freelancers with Generative AI researchers and product teams working on cutting-edge AI systems. We partner with independent contractors to support large-scale model evaluation, alignment, and quality initiatives.

As a Freelance Academic Research Expert, you will pair genuine scholarly expertise with hands-on GenAI-evaluation experience to design challenging, realistic prompts backed by source documents such as peer-reviewed papers, published datasets, methodology standards, and systematic review protocols. You will then probe where the model breaks, and encode expert judgment into precise, gradable rubrics and a canonical 'golden' answer. You own an end-to-end evaluation task from prompt design through reviewer sign-off.

This is hands-on academic work with an analytical core. You will produce scholarly deliverables as the output; what makes the work difficult is evaluating methodology, evidence, and citation integrity against the provided sources. It is not a teaching, curriculum delivery, or tutoring role.

 

What you'll work on

  • Apply your academic research expertise to help create and evaluate complex AI prompts and responses.
  • Conceptualize and draft realistic, multi-step prompts that mirror real research assignments and require synthesis, methodological evaluation, and reconciliation across multiple sources.
  • Identify and document genuine model failures by writing specific, objective, and actionable explanations.
  • Build comprehensive evaluation rubrics that trace each required input through dependent criteria and carefully separate extraction from interpretation.
  • Produce fully correct, deliverable-quality golden responses that satisfy all criteria using discipline-standard conventions and citation practice.
  • Complete our onboarding process and strictly follow the disciplined authoring and review workflow for each task.
  • Closely follow guidelines, run AI quality checks on your own work, and resolve any issues in the feedback loop before final submission.

Engagement details

  • Location: Remote (United States)
  • Engagement Type: 1099 Freelance / Independent Contractor, through the Uber AI Solutions platform
  • Engagement Duration: 40 hours over 1-2 weeks

Who we're looking for

  • Ph.D., Ed.D., or completed Postdoctoral fellowship from an accredited institution.
  • 5+ years of experience in academic research, peer review, or higher education instruction.
  • Demonstrated ability to read primary source documents (peer-reviewed literature, published datasets, methodology standards) and derive multi-step analytical conclusions from them.
  • Familiar with research methodology evaluation, systematic literature review, and citation integrity standards.
  • Comfortable working independently on detail-oriented tasks.
  • Strong analytical and written communication skills.
  • Prior hands-on experience with GenAI / LLM evaluation and prompt engineering, rubric or golden-answer authoring, data annotation, RLHF, red-teaming, or model quality assessment.

 

Why this matters

Your expertise will help improve how AI systems handle complex scholarly topics, making outputs more accurate, reliable, and better aligned with real-world research reasoning.