Member of Technical Staff — Training Infrastructure

Posted yesterday

human intuitionNew York (NY)

SENIORITY

Mid

Apply

About the role

Building the autonomous company Human Intuition is building the autonomous company. Businesses run on accumulated judgment: how to interpret a situation, choose an action, and learn from its consequences. Much of that knowledge lives in people, even when the decisions they make leave traces in software. We are working to make that judgment learnable. A business has defined systems, tools, permissions, histories, and objectives. Those boundaries create an opportunity to build agents that learn from how work is done, act within clear constraints, and improve through feedback. Our ambition is to turn the knowledge inside institutions into software that compounds.
The role: Build the platform that turns research ideas into reproducible learning runs. You will connect datasets, environments, rollout generation, trainers, checkpoints, and evaluations into a system researchers can understand and operate. The aim is to shorten the path from a question to a trustworthy result.
What you’ll do:
  • Build orchestration for fine-tuning and reinforcement learning workloads across the compute resources the team uses.
  • Coordinate training, inference, and environment workers with clear job state and failure recovery.
  • Preserve experiment lineage across data, configuration, code, model versions, and evaluation results.
  • Improve checkpointing, artifact storage, job resumption, and resource utilization.
  • Provide useful logs, metrics, and debugging tools for learning and infrastructure failures.
  • Work with researchers to make new methods repeatable and with engineers to deliver validated models to serving systems.
  • What you’ll bring
  • Experience with machine learning infrastructure or distributed systems used for compute-intensive work.
  • Strong Python and practical familiarity with training workloads and accelerator constraints.
  • Experience with workload orchestration, cloud infrastructure, or container platforms.
  • An ability to debug across application code, workers, networking, storage, and resource management.
  • Care for reproducibility, operational simplicity, and the experience of the people using the platform.
  • Useful experience
Distributed training, rollout systems, experiment tracking, model registries, GPU scheduling, or post-training infrastructure.
What success looks like: Researchers can launch, inspect, reproduce, and recover experiments with confidence, and successful results move into deployment with a clear record of how they were produced.

Before you apply

Applying takes about a minute. These four things decide how fast it moves after that.

Your profile is current

It's what we read first. Occupations, seniority and locations matter more than a long history.

Two examples you can talk through

Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.

A number in mind

What you're on now and what would make you move. We negotiate better when we know both.

Your notice period

Employers plan around it, and it's the question that stalls offers most often.

Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.

More like this