About the role

Directly focused on building benchmarks and evaluation frameworks for LLMs and AI models — closely tied to vibe-coding and benchmarking work.
About the Role: Lead Vals AI's research efforts to design and validate new evaluation methodologies and benchmarks for LLMs and other AI systems, set research direction, publish influential work, and build and manage a research team while working closely with enterprise customers and lab partners in San Francisco.
Job Description: Role Head of Research responsible for defining and advancing evaluation methodologies and benchmarks for large language models and other AI systems. The role sets research direction across Vals’ portfolio, publishes work that moves the field forward, recruits and grows a research team, and partners directly with enterprise customers and AI labs on real‑world evaluation problems.
Key Responsibilities: Develop new paradigms for evaluating long-horizon, real‑world tasks that current benchmarks and judge‑model approaches fail to capture. Oversee and set direction across Vals’ research portfolio and ongoing projects. Publish research and present results to the community, customers, and partners. Recruit, mentor, and grow a high-quality research team alongside the founders. Collaborate closely with enterprise customers and lab partners to solve practical evaluation problems.
Requirements: PhD in ML/NLP (in progress or completed) or equivalent industry research track record. Deep familiarity with the LLM evaluation landscape, including existing benchmarks, failure modes, judge‑model approaches, and human‑in‑the‑loop methodologies. Preference for research that influences real‑world deployments rather than easily gamed benchmarks. Strong written and verbal communication skills for publishing, presenting, and customer engagement. Ability to work in‑person in San Francisco.
Nice‑to‑Haves: A widely cited benchmark or evaluation framework you built or co‑built. Prior experience at a frontier lab (e.g., Anthropic, OpenAI, Google DeepMind, Meta FAIR) or a research‑led startup. Domain depth in verticals such as legal, finance, insurance, or healthcare. Experience leading or mentoring other researchers and maintaining a public research presence (papers, blog posts, talks, OSS contributions).What We OfferHighly competitive salary and equity. Relocation and transportation support. Health and dental insurance coverage. Lunch and dinner provided, plus free snacks, coffee, and drinks.401(k) plan. Unlimited PTO.$1,500 housing stipend (within one‑mile radius).Collaboration with leading AI labs and opportunity to work on foundational evaluation research. PythonDjangoReactAWSAWS CDKCompensation
Compensation Range: $225,000 - $275,000 (USD)
Skills Research Leadership Experimental Design LLM Evaluation Publication & Communication Team Building Mentorship Customer‑facing Collaboration Domain Expertise Learning Agility Ownership Problem Solving Experience Level USD 225,000 - 275,000/year Employment Type Full-time Relocation and transportation support Dental insurance Lunch and dinner provided Free snacks/coffee/drinks$1,500 housing stipend (within one‑mile radius)

Before you apply

Applying takes about a minute. These four things decide how fast it moves after that.

Your profile is current

It's what we read first. Occupations, seniority and locations matter more than a long history.

Two examples you can talk through

Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.

A number in mind

What you're on now and what would make you move. We negotiate better when we know both.

Your notice period

Employers plan around it, and it's the question that stalls offers most often.

Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.

More like this