Job Description
About This Opportunity
We are currently building a talent pool of experienced AI Evaluation Engineers for an upcoming client project focused on AI model assessment, benchmarking, and quality assurance.
This is a freelance, project-based engagement, with project scope, duration, workload, and onboarding timelines currently being finalized. We welcome expressions of interest from qualified professionals with relevant technical and evaluation experience.
At this stage, we are proactively connecting with qualified professionals who may be interested in participating in the project. Candidates who are a potential fit will be considered and contacted as the project details are confirmed and the engagement moves forward.
Job Summary
We are seeking an AI Evaluation Engineer to assess, validate, and improve AI task quality across the end-to-end evaluation pipeline. The ideal candidate possesses strong software engineering skills, hands-on experience with Python-based environments, and the ability to identify root causes of issues within AI evaluation, benchmarking, and rollout processes.
Key Responsibilities
- Evaluate and analyze AI tasks, benchmarks, and execution traces to identify quality issues and performance gaps across the pipeline.
- Investigate root causes of failures and determine whether issues originate from task design, evaluation metrics, reference solutions, execution environments, or rollout contamination.
- Validate evaluation methodologies, ensuring reproducibility, environment consistency, model discrimination, and proper visible/hidden evaluation isolation.
- Execute and troubleshoot tasks locally using Python, Shell, Docker, and AI coding tools such as Claude Code or Codex.
- Review evaluation results, artifacts, logs, and traces, providing evidence-based recommendations and quality assessments.
- Produce clear Pass, Conditional Pass, or Reject decisions with actionable improvement plans and comprehensive documentation.
Requirements
- Bachelor's degree in a relevant technical field or equivalent practical experience.
- Strong hands-on programming experience in Python, Shell scripting, and Docker-based environments.
- Experience in AI model evaluation, benchmarking, testing, validation, quality assurance, or research operations.
- Ability to analyze execution traces, debugging logs, evaluation outputs, and identify root causes across complex workflows.
- Hands-on experience using AI coding tools such as Claude Code, Codex, or similar tools to test, debug, and validate tasks locally.
- Strong analytical thinking, problem-solving, and technical reporting skills.
Preferred Qualifications
- Experience in AI/ML research, model evaluation, reinforcement learning, or benchmark development.
- Experience designing evaluation methodologies and ensuring reproducible experiment results.
- Knowledge of model comparison, testing frameworks, and AI quality assessment practices.
What Success Looks Like
- Consistently identify and resolve quality issues within AI evaluation pipelines.
- Deliver accurate, evidence-backed evaluation conclusions.
- Improve benchmark quality, reproducibility, and model assessment reliability.
- Support the development of high-quality datasets and evaluation standards for AI systems.

