Remote | Senior Software Engineer - AI Evaluation / Coding Agents $100-$150/hour

Posted today

24magDenver (CO)

SENIORITY

Senior

SALARY

$100.00 - $150.00 per hour

Apply

About the role

Senior Software Engineer - AI Evaluation / Coding AgentsWe are sharing a specialised freelance opportunity for experienced software engineers to evaluate and improve advanced AI coding systems through rigorous code review, repository-based testing, rubric development, and engineering-quality assessment. Selected professionals will work with coding agents across substantial real-world codebases, review generated implementations, identify technical failure modes, and translate expert engineering judgement into structured evaluation signals. The focus is not primarily on building production applications, but on determining whether AI-generated software is correct, robust, maintainable, and aligned with professional engineering standards.
Key Responsibilities:
  • AI-Generated Code Evaluation
  • Review code produced by AI coding agents
  • Assess implementations for correctness, robustness, and maintainability
  • Determine whether agents selected appropriate technical approaches
  • Identify subtle implementation errors and engineering weaknesses
  • Explain clearly why generated solutions succeed or fail
  • Rubric & Preference Evaluation
  • Design and refine technical evaluation rubrics
  • Define criteria for assessing coding-agent performance
  • Review and label preference and evaluation data
  • Apply consistent quality standards across repeated assessments
  • Capture nuanced differences between competing implementations
  • Repository & Engineering Analysis
  • Work within substantial real-world software repositories
  • Analyse code changes in their broader architectural context
  • Evaluate repository-level behaviour rather than isolated snippets
  • Review implementation quality using senior engineering judgement
  • Identify recurring failure patterns across coding tasks
  • Evaluation Infrastructure & Workflows
  • Build and maintain pipelines supporting data generation and evaluation
  • Improve infrastructure for collecting and reviewing model outputs
  • Support scalable evaluation and feedback workflows
  • Contribute to reliable processes for repeated technical assessment
  • Help translate qualitative engineering judgement into structured systems
  • Research & Technical Collaboration
  • Collaborate with research and engineering teams
  • Summarise evaluation findings in clear written reports
  • Provide actionable recommendations based on observed model behaviour
  • Help improve evaluation methodologies and coding-agent benchmarks
  • Communicate complex technical issues in concise, structured language
  • Ideal Profile 5+ years of hands-on software engineering experience
  • Strong proficiency in Python, TypeScript, JavaScript, Go, Java, or another major production language
  • Experience working in substantial real-world codebases
  • Strong code-review and technical-analysis skills
  • Ability to assess implementation correctness and maintainability
  • Strong understanding of modern software-engineering practices
  • Experience with GitHub-based development workflows
  • Familiarity with CI/CD and production engineering processes
  • Comfortable analysing unfamiliar repositories and code changes
  • Ability to identify subtle technical issues others may overlook
  • Strong written communication and structured technical reasoning
  • Experience using modern LLMs or AI-assisted coding tools
  • Familiarity with open-source development is valuableLLM evaluation or coding-agent experience is advantageous
  • Experience with RLHF, preference data, rubric design, or post-training is useful but not required
  • Engagement Details
  • Independent contractor engagement
Remote: North America only
Compensation: $100–$150/hour
Candidates should provide a specific hourly rate expectation Flexible commitment of approximately 20–40 hours per week At least 6 hours of Pacific Time overlap per day is required Expected project duration is approximately 3 months Start is as soon as possible Evaluation process includes an approximately 25-minute AI interview A practical code and AI-evaluation exercise of approximately 30 minutes follows Final stage includes an approximately 20-minute hiring manager interview The practical exercise focuses on reviewing AI-generated code rather than competitive programming or algorithm puzzles Work must be completed without using confidential, proprietary, unreleased, employer-restricted, client-restricted, or otherwise protected code, repositories, datasets, architecture materials, or technical information belonging to any employer, client, institution, or other third party This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy

Before you apply

Applying takes about a minute. These four things decide how fast it moves after that.

Your profile is current

It's what we read first. Occupations, seniority and locations matter more than a long history.

Two examples you can talk through

Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.

A number in mind

What you're on now and what would make you move. We negotiate better when we know both.

Your notice period

Employers plan around it, and it's the question that stalls offers most often.

Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.

More like this