AI Evals Engineer - Evaluation Datasets & Ground Truth
AI Evals Engineer - Evaluation Datasets & Ground Truth
Posted 2 days ago
SENIORITY
Lead
About the role
- Decide what needs to be measured, and how
- Read our pipelines and system architecture, sit with product and engineering, and decompose each system into evaluable modules with explicit input → expected-output contracts.
- For each module, define what "correct" means in writing: rubrics, label schemas, edge-case policies, and the tolerances that matter to the business.
- Prioritize. We have more modules than you can cover in year one; you'll decide where a validation set unblocks the most iteration, with minimal direction from principal engineers.
- Source the data by whatever means is cheapest and most trustworthy for that module
- Production sampling: pull stratified, de-identified samples from real traffic so eval sets reflect what the system actually sees, including the long tail.
- Human labeling: scope and run labeling programs, write annotation guidelines, build calibration sets, measure inter-annotator agreement, and manage vendors (including offshore labeling teams) or internal subject-matter experts. You own label quality, not just label throughput.
- Must have
- Experience building or evaluating ML or LLM-powered systems in production, in a role where output quality was your problem.
- You have built evaluation or validation datasets before and can talk about one in detail: how you defined correctness, how you sourced labels, what went wrong, how you knew the labels were good.
- Working fluency in ML validation fundamentals: train/validation/test discipline, stratified sampling, precision/recall trade-offs, class imbalance, calibration, inter-rater agreement, basic significance testing and power.
- Understand feature engineering well enough to reason about why a classifier fails and what data would expose it.
- Strong Python and SQL; comfortable pulling and reshaping data yourself.
- Hands-on with LLM-based systems: prompting, structured outputs, agent/tool-use harnesses, and the specific ways they fail (non-determinism, prompt sensitivity, evaluator bias).Judgment about when LLM-as-judge is reliable and when it is not, and how to prove either.
- You can read a system design, understand the business logic it encodes, and translate that into a label schema without waiting to be told.
- Have run a human-labeling program end to end, including vendor selection, guideline authoring, QA sampling, and cost/quality trade-offs.
- Experience with eval tooling.
- Experience with data labeling platforms.
- Have used a frontier model as a distillation/oracle source and can articulate where that assumption breaks.
- How you think
- Verify the verifier. A label is a claim, not a fact, until something independent agrees with it.
- Start simple: binary before graded, one judge before five, a hundred well-understood examples before ten thousand noisy ones.
- You measure business outcomes, not model vibes, and you're comfortable telling a senior engineer their favorite change didn't move the number.
- You'd rather own an unglamorous problem completely than a glamorous one partially.
- Team & reporting
- You'll report to the Chief AI Officer and work across all product/pipeline teams. Over time it is expected that you'll lead a small team of eval engineers. You'll have a generous budget for labeling vendors and oracle compute.
- Why Prophetic:
- The team. Wide expertise, sharp, low ego, and genuinely fun to work with. We have each other's backs and we celebrate wins together.
- Real ownership and impact from day one. Our customers make million-dollar decisions on our platform. Your work matters immediately.
- AI-native product and AI-native workflows. Every team - engineering, sales, operations - runs on AI tooling daily.
- Founder-led transparency. Leadership shares the financials, the strategy, and the hard calls with the whole team.
- High growth, high demand. The product works, customers love it, and we are hiring to keep up. No shortage of opportunity here.
- Flexibility. Remote, hybrid, or in-office with a great Portland, OR headquarters.
- Self-starters who thrive in ambiguity. We're a startup. If you need someone to tell you what to do every morning, this isn't the place; if you come alive solving hard problems with smart people, it is.
- (US-based Full-time):100% medical, dental & vision insurance coverage for you; 30% coverage for dependents
- Competitive salary and meaningful early-stage equity
- Unlimited PTOHybrid/Remote stipend
Before you apply
Applying takes about a minute. These four things decide how fast it moves after that.
Your profile is current
It's what we read first. Occupations, seniority and locations matter more than a long history.
Two examples you can talk through
Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.
A number in mind
What you're on now and what would make you move. We negotiate better when we know both.
Your notice period
Employers plan around it, and it's the question that stalls offers most often.
Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.
More like this
