AI Data Engineer

Posted today

howard hughes medical instituteChevy Chase (MD)

SENIORITY

Lead

SALARY

$128,817 per year

Apply

About the role

The Howard Hughes Medical Institute (HHMI) is building an AI Accelerator, and this role helps power the foundation behind it. As an AI Data Engineer, you will design, build, and operate governed, AI-ready data pipelines that turn institutional content into data products for downstream retrieval and AI developer use. Work is performed on HHMI’s Databricks-based platform under the design authority of the Principal Knowledge & Data Architect. This position is based inChevy Chase, MDand workshybrid, withthree days per week in-personat HHMI offices.
What you will do:
  • Build AI-facing data pipelines, including ingestion from Operations Capabilities landing zone inputs, transformation throughraw bronze , silver, andgoldlayers, and serving ofgoverned, AI-readycontent for downstream consumption.
  • Implement themedallion architectureby turning K&D Architect design patterns into working pipelines, tables, and materialization schedules.
  • Build and operate
  • Delta Laketables with partitioning, optimization, and evolution in mind.
  • Own workflow orchestration using
  • Databricks Workflowsand
  • Delta Live Tables, includingretry policies, alerting routes, run history, andcost tags .Implement governance patterns using
  • Unity Catalog, including sensitivity classification, access control, and audit for AI-facing data assets.
  • Build and operate retrieval-supporting infrastructure, including embedding pipelines, vector store maintenance, reindexing when models upgrade, and retrieval evaluation frameworks.
  • Partner with Operations Capabilities to definesource-system contractsfor what the AI Fabric consumes at the landing zone, including schema, cadence, SLA, and quality thresholds.
  • Partner with Operations Capabilities to own theplatform-sideof that contract.
  • Design and operate data quality and observability: data-quality checks, freshness monitoring, drift detection, and alerting that surfaces issues before they reach AI users.
  • Support AI Developer velocity by delivering the data layer AI Developers consume when use cases require specific data.
  • Contribute to and consume the sharedreference-pattern library, including reusable pipeline patterns, code templates, and platform standards.
Key requirements:
  • At least 4 yearsof hands-on production data engineering experience designing, building, and operating production data pipelines.
  • Deep experience withDatabricks and Spark, including
  • Delta Lake , medallion architecture , Delta Live Tables , Databricks Workflows , Databricks SQL, and
  • Unity Catalog .Python and SQL fluency, including PySpark and SQL that runs at scale, plus ETL patterns.
  • Working knowledge ofGit , CI/CD for data pipelines, andinfrastructure-as-codewith
  • Terraform .AI-adjacent data engineering experience building foundations for AI use cases, including embedding pipelines, vector stores, chunking strategies, and retrieval evaluation.
  • Workflow orchestration experience withDatabricks Workflowsor
  • Airflowin production, including retry semantics, dependency management, and failure handling.
  • Data quality and observability experience with
  • Great Expectations , Databricks data-quality monitors, or equivalent.
  • Governance discipline with
  • Unity Catalogstructures and sensitivity classification, designing for access control and audit from the first commit.
  • AWS foundations at the level needed for Databricks-on-AWS work:
  • IAM , S3, andKMS .Communication skills working with the K&D Architect, AI Developers, and Operations Capabilities, including explaining data-engineering trade-offs to non-engineers.
  • Bachelor’s degree or equivalent, plus meaningful exposure to AI or knowledge-management use cases alongside the required data engineering experience.
Technologies Databricks Workflows, Delta Live Tables, Delta Lake, Databricks SQL, Unity Catalog Python, PySpark, SQL, GitCI/CD, Terraform Databricks data-quality monitors, Great Expectations Embedding pipelines, vector store maintenance, retrieval evaluation frameworks Databricks-on-AWS, IAM, S3, KMSAirflowCompensation and benefits Hiring pay range:$128,816.80 - $161,021.00 (annual) Competitive pay Exceptional health benefits Retirement plans Time off Range of recognition and wellness programs Reporting and employment details Reports to the Director of AI Enablement .HHMI is not able to sponsor a visa for this position at this time. HHMI usesE-Verifyto confirm identity and employment eligibility for new hires.
Additional information: The description outlines principal duties and responsibilities and is not necessarily exhaustive. Unless they begin with the word “may,” described essential duties and responsibilities are essential functions under the Americans with Disabilities Act.

Before you apply

Applying takes about a minute. These four things decide how fast it moves after that.

Your profile is current

It's what we read first. Occupations, seniority and locations matter more than a long history.

Two examples you can talk through

Not a portfolio — just two pieces of work where you can explain the decisions and what you'd change.

A number in mind

What you're on now and what would make you move. We negotiate better when we know both.

Your notice period

Employers plan around it, and it's the question that stalls offers most often.

Once you apply, someone reads it and calls you before anything reaches the employer — usually within two working days.

More like this