Job Description
Our Mission
Reflection is a research lab making intelligence open and accessible for everyone to use, customize, and build on. We build open models that let anyone control their intelligence and help shape the future of AI. Our mission: make intelligence open and accessible to all.
About the Role
Conduct critical comparative analysis to advance our understanding of model capabilities
Build and refine evaluation systems and processes that create tight feedback loops between data, evals, and model behavior
Develop generalizable evaluation frameworks that capture what matters for reasoning, alignment, and usefulness.
Collaborate closely with pre-training, post-training, and applied teams to translate insights into model improvements.
Push the boundaries of what’s measurable, from synthetic evals to human feedback and real-world interaction data.
About You
Strong statistical analysis and experimental design skills to rigorously measure model improvements
Familiarity with LLM evaluation methodologies: static benchmarks, human preference evals, and/or agentic tasks.
High agency and thrive in a fast-paced startup environment; bias for impact over process.
Excited to work in a new frontier lab, defining how we measure and accelerate progress toward more capable models.
Collaborative, detail-oriented, and motivated by building the feedback loops that make models truly improve.
What We Offer:
We believe that to make intelligence open and accessible to all, you need to start at the foundation. Joining Reflection means building from the ground up as part of a talent-dense team. You will help define our future as a company, and help define the future of open foundational models.
We want you to do the most impactful work of your career with the confidence that you and the people you care about most are supported.
Top-tier compensation: Salary and equity structured to recognize and retain our talent globally.
Stock options: Everyone who joins and contributes to Reflection's success gets to share in the upside through stock options.
Health & wellness: Comprehensive medical, dental, vision, and life, with an annual wellness allowance.
Meals: Lunch and dinner are provided in the office daily.
Life & family: 22 weeks paid parental leave for all new birthing and non-birthing parents, including adoptive and surrogate journeys.
Vacation days: Unlimited paid time off in the U.S. and 30 days in the U.K.
Sponsorship support: We sponsor visas to help exceptional talent join our team and support long-term immigration pathways where applicable.
Team building: We have regular off-sites, happy hours, and team celebrations.
Categories
Frequently asked questions
Is the Member of Technical Staff - Evaluations position at Reflection AI remote?
The Member of Technical Staff - Evaluations role at Reflection AI is an on-site or hybrid position.
What type of employment is the Member of Technical Staff - Evaluations role?
Reflection AI is hiring for a full-time Member of Technical Staff - Evaluations position.
How do I apply for the Member of Technical Staff - Evaluations position at Reflection AI?
You can apply for the Member of Technical Staff - Evaluations role directly through Reflection AI's official application link provided on this page.
Similar AI jobs
Sr. Integration & Test Engineer - Effector Systems
Chaos Industries · fulltime
Chip Design Engineer
NVIDIA · fulltime
Chip Design Verification Engineer
NVIDIA · fulltime
Part-time Student Worker – Software Development Engineer in Test
Zoox · contract
Spectrum Sensing Engineer
Chaos Industries · fulltime
Part-Time Student Worker – AI Validation and Benchmarking Engineer
Zoox · contract