Join a frontier AI research lab developing the next generation of foundation models. Help design the evaluations, experiments, and benchmarks that define how state-of-the-art AI systems are measured and improved.
Overview
A frontier AI research lab is expanding its Model Evaluation team and is looking for Machine Learning Engineers with strong research and experimentation experience. In this role, you'll build challenging machine learning tasks that test the capabilities of today's most advanced AI models across implementation, experimentation, and technical reasoning.
You'll design realistic ML workflows, implement reference solutions, run experiments, and analyze results to establish high-quality ground truth for model evaluation. Your work will directly influence how frontier models are benchmarked, improved, and deployed.
Working alongside research scientists and ML engineers, you'll contribute to the evaluation pipeline for cutting-edge generative AI systems while solving technically demanding problems drawn from real machine learning practice.
What You'll Do
Design evaluation tasks: Create realistic machine learning workflows that assess the implementation and reasoning capabilities of frontier AI models.
Build reference solutions: Develop experiments in Python, train and evaluate models, and produce technically rigorous benchmark solutions.
Evaluate model performance: Review AI-generated implementations, identify failure modes, and document technical reasoning behind model successes and shortcomings.
Develop evaluation assets: Create benchmark datasets, scoring rubrics, and validation criteria used to measure model quality across diverse ML tasks.
Collaborate with researchers: Partner with research scientists and fellow engineers to improve benchmark quality, consistency, and scientific rigor.
What We're Looking For
- Master's degree or PhD in Machine Learning, Computer Science, Artificial Intelligence, or a related quantitative discipline.
- Experience conducting machine learning research or building production-quality ML systems.
- Strong hands-on experience designing experiments, training models, and interpreting experimental results.
- Proficiency in Python, Git, Jupyter, and modern deep learning frameworks such as PyTorch or TensorFlow.
- Familiarity with large language models, AI evaluation methodologies, or benchmark development is preferred.
- Experience with reinforcement learning, optimization, or model alignment is a plus.
- Excellent technical writing skills, strong attention to detail, and the ability to solve open-ended engineering problems independently.
- Availability to work approximately 15-20 hours per week.
Contract and Payment Terms
- You'll be engaged as an independent contractor; projects can be extended, shortened, or concluded based on need and performance.
- Fully remote, completed on your own schedule.
- Rates are shown upfront per project.
- Payment is made on a per-project or scheduled basis for work delivered.