Senior Machine Learning Engineer

Posted yesterday

medici land governanceDenver (CO)

SENIORITY

Manager

Apply

About the role

Lead ML Engineer, Document Intelligence Platform We’re hiring a hands-on technical leader to own our end-to-end document intelligence platform. Our system processes court documents, property records, and other complex scanned files across both clean new recordings and poor-quality historical documents. We need someone who can lead architecture and execution across document quality assessment, preprocessing, OCR/VLM extraction, indexing, provider fallback, evals, and serving strategy. This is a high-ownership role for someone who can make strong technical decisions, improve production accuracy, reduce hallucinations, and build resilient systems in a fast-changing model landscape. What you’ll owndocument intake, quality assessment, and routingimage preprocessing and scan enhancementOCR, VLM, and document extraction workflowsstructured extraction and indexingfallback across multiple AI providers and standalone OCR/model systemsevals, regression testing, and prompt/model versioningserving strategy for batch and real-time workloadslatency, throughput, and memory optimizationconfidence scoring, validation, and hallucination reduction
What you’ll do:
  • Architect and improve document AI pipelines for court records, property records, and other difficult document sets
  • Build systems that assess image quality and document complexity on arrival and determine the correct processing path
  • Optimize preprocessing for skew correction, denoising, deblurring, contrast enhancement, binarization, orientation correction, cropping, and scan recovery
  • Design robust fallback across multiple paid model providers and standalone OCR/VLM systems to reduce vendor dependency
  • Evaluate platforms such as direct APIs, OpenRouter, Vertex AI, and self-hosted inference for resiliency, cost, and operational flexibility
  • Assess paid API vs self-hosted economics, including infrastructure cost, memory usage, throughput, fine-tuning needs, and long-term maintainability
  • Build or optimize self-hosted serving systems using vLLM or comparable inference stacks
  • Support both offline batch workloads and lower-latency live endpoints, including queueing and prioritization tradeoffs
  • Improve compatibility between OCR outputs, custom internal formats, and downstream indexing/extraction systems
  • Build an evals framework using MLflow or equivalent tooling, including ground-truth datasets, regression testing, prompt versioning, and user-feedback integration
  • Establish measurable standards for OCR accuracy, extraction accuracy, hallucination rate, indexing quality, latency, throughput, and cost per document
  • Lead failure analysis and iterative improvement, especially for degraded historical records
  • Provide technical leadership to a small AI engineering team while remaining deeply hands-onWhat we’re looking for 6+ years building production AI/ML systems
  • Strong experience in document AI, OCR, computer vision, NLP, or information extraction
  • Experience owning or materially improving intelligent document processing pipelines
  • Experience evaluating multiple model providers, APIs, and document AI services in production
  • Strong understanding of paid APIs vs open-source/self-hosted tradeoffs
  • Hands-on experience with vLLM, Triton, TGI, or similar inference-serving systems
  • Experience with tools such as Tesseract, PaddleOCR, Textract, Azure Document Intelligence, Google Document AI, LayoutLM, Donut, or similar
  • Strong Python skills and experience with PyTorch, OpenCV, Hugging Face, and modern ML infra
  • Experience building eval frameworks with regression testing, ground truth, versioning, and production monitoring
  • Strong understanding of batching, scheduling, GPU memory, latency, throughput, and cost optimization
  • Experience reducing hallucinations through validation, confidence scoring, schema enforcement, and guardrails
  • Ability to lead technical direction and drive execution in an ambiguous environment
  • Strong bonus points
  • Court records, property records, legal documents, or public-records experience
  • Historical document digitization or degraded scan handling
  • Hybrid OCR + VLM + rules + human-review system designML/LLM observability using MLflow, Langfuse, W&B, or similar
  • Priority scheduling, queue design, and mixed batch/live inference systems
  • Fine-tuning document models when vendor APIs are insufficient
  • Experience leading a small team or owning a critical AI platform
  • Who thrives here
  • Someone who wants real ownership
  • Someone who can make decisions and move
  • Someone who is hands onSomeone who likes messy data and operational complexity
  • Someone who cares about accuracy, resiliency, speed, and cost
  • Someone who can lead without hiding behind process
  • Who is not a fit
  • Someone looking for a pure research role
  • Someone who only wants prompt engineering or prototyping
  • Someone who has not owned production systems with reliability and cost constraints
  • Someone who prefers narrow model work over full-platform responsibility

Before you apply

Applying takes about a minute. These four things decide how fast it moves after that.

Your profile is current

It's what we read first. Occupations, seniority and locations matter more than a long history.

Two examples you can talk through

Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.

A number in mind

What you're on now and what would make you move. We negotiate better when we know both.

Your notice period

Employers plan around it, and it's the question that stalls offers most often.

Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.

More like this