Talent.com
Senior Data Engineer II - Data Curation

Senior Data Engineer II - Data Curation

Formation BioNew York, NY, United States
3 days ago
Job type
  • Full-time
Job description

About Formation Bio

Formation Bio is a tech and AI driven pharma company differentiated by radically more efficient drug development.

Advancements in AI and drug discovery are creating more candidate drugs than the industry can progress because of the high cost and time of clinical trials. Recognizing that this development bottleneck may ultimately limit the number of new medicines that can reach patients, Formation Bio, founded in 2016 as TrialSpark Inc., has built technology platforms, processes, and capabilities to accelerate all aspects of drug development and clinical trials. Formation Bio partners, acquires, or in-licenses drugs from pharma companies, research organizations, and biotechs to develop programs past clinical proof of concept and beyond, ultimately helping to bring new medicines to patients. The company is backed by investors across pharma and tech, including a16z, Sequoia, Sanofi, Thrive Capital, Sam Altman, John Doerr, Spark Capital, SV Angel Growth, and others.

You can read more at the following links :

  • Our Vision for AI in Pharma
  • Our Current Drug Portfolio
  • Our Technology & Platform

At Formation Bio, our values are the driving force behind our mission to revolutionize the pharma industry. Every team and individual at the company shares these same values, and every team and individual plays a key part in our mission to bring new treatments to patients faster and more efficiently.

About the Position

Formation Bio is seeking a hands-on technical leader to shape the future of our Data Curation team. This role is ideal for someone passionate about modeling, harmonizing, and unifying complex biomedical and healthcare data into high-quality, stakeholder-ready assets. You'll lead efforts to transform structured, semi-structured, and unstructured data into trusted, interoperable models that power analytics, product development, and scientific decision-making across the company.

This is a high-impact role for a builder who thrives on technical depth, ontology-driven modeling, and architectural leadership. You'll set standards for how Formation Bio integrates healthcare and life sciences ontologies across diverse datasets, and you'll guide the design of the technical stack that enables interoperability, reuse, and long-term usability of curated data assets. You'll partner closely with domain experts in areas like claims, EHR, and biomedical research - ensuring their deep knowledge is translated into consistent, governed, and reusable data products.

Responsibilities

Technical Leadership & Strategy

  • Define and communicate technical direction for the Data Curation team.
  • Drive the architecture and technical stack for ontology-driven harmonization across healthcare and pharmaceutical datasets.
  • Partner with domain experts (claims, EHR, pharma, research) to align technical standards across diverse datasets.
  • Mentor engineers in best practices for modeling, ontology integration, and scalable curation workflows.
  • Data Modeling, Ontology-Driven Harmonization & Unstructured Integration

  • Lead development of robust SQL / dbt models that unify complex healthcare and pharma datasets.
  • Apply healthcare and biomedical ontologies (e.g., SNOMED, RxNorm, UMLS, Mondo, OMOP, FHIR) to ensure interoperability and consistent integration.
  • Design scalable workflows for ontology alignment, normalization, and harmonized data product creation .
  • Lead integration of unstructured data sources (clinical notes, publications, documents, scientific text) using NER, NLP, embeddings, and document parsing .
  • Define patterns for linking structured and unstructured assets into a unified semantic layer that is easy to query, search, and consume.
  • Establish standards for vector database usage and semantic search , ensuring embeddings and structured models are connected and interoperable.
  • Knowledge Integration & Architecture

  • Establish architectural patterns for managing ontology mappings, ontology-driven transformations, and harmonized knowledge assets .
  • Define how graph-based and relational representations complement each other for interoperability.
  • Collaborate with the Data Infrastructure team to align ingestion, orchestration, and governance frameworks with curation needs.
  • Data Quality & Catalog Stewardship

  • Enforce structural and semantic quality standards with automated checks and validation.
  • Maintain and enrich the enterprise data catalog, ensuring curated datasets - structured and unstructured - are discoverable and well-documented.
  • Capture and codify domain-specific knowledge into durable, governed data assets.
  • About You

  • 7+ years of experience in data engineering, semantic modeling, or data curation, with leadership experience in technical direction.
  • Proven expertise in SQL / dbt modeling and integrating healthcare and biomedical ontologies .
  • Hands-on experience with ontology-driven harmonization and data model integration across heterogeneous datasets.
  • Strong background in data architecture and stack design , with the ability to define standards and paved paths.
  • Experience working with unstructured data : entity extraction (NER), NLP, embeddings, or document parsing.
  • Familiarity with vector databases, semantic search, and knowledge graph concepts - and how to connect these with structured datasets for unified consumption.
  • Comfortable with Python, orchestration tools (Dagster, Airflow), and working with diverse data types.
  • Skilled at collaborating with infrastructure teams to balance semantic integration with scalable foundational tooling.
  • Excited to mentor others, set high standards, and drive alignment across a multidisciplinary team.
  • Bonus points if you have :

  • Experience with knowledge graph technologies (e.g., Neo4j, RDF / SPARQL, Cypher).
  • Experience with healthcare and life sciences ontologies such as Mondo, OMOP, FHIR, SNOMED, RxNorm, UMLS .
  • Experience harmonizing datasets from EHR, claims, or biomedical research domains.
  • Contributions to enterprise data catalogs or metadata management frameworks.
  • Formation Bio is prioritizing hiring in key hubs, primarily the New York City and Boston metro areas, with additional growth in the Research Triangle (NC) and San Francisco Bay Area. Please only apply if you reside in these locations or are willing to relocate.

    Compensation :

    The target salary range for this role is : $220,000 - $280,000.

    Salary ranges are informed by a number of factors including geographic location. The range provided includes base salary only. In addition to base salary, we offer equity, comprehensive benefits, generous perks, hybrid flexibility, and more. If this range doesn't match your expectations, please still apply because we may have something else for you.

    You will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, or veteran status.

    #LI-hybrid

    Create a job alert for this search

    Data Engineer Ii • New York, NY, United States

    Related jobs
    • Promoted
    Senior Data Engineer II

    Senior Data Engineer II

    VirtualVocationsJersey City, New Jersey, United States
    Full-time
    A company is looking for a Senior Data Engineer II to join their data engineering team.Key Responsibilities Design, develop, and maintain scalable data pipelines using Apache Spark on Databricks ...Show moreLast updated: 1 day ago
    • Promoted
    Senior Data Engineer

    Senior Data Engineer

    VirtualVocationsBronx, New York, United States
    Full-time
    A company is looking for a Senior Data Engineer to lead the data engineering function and enhance data capabilities.Key Responsibilities Design, build, and optimize data architecture, pipelines, ...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Data Engineer II - Data Curation

    Senior Data Engineer II - Data Curation

    BioSpace, Inc.New York, NY, United States
    Full-time
    Formation Bio is a tech and AI driven pharma company differentiated by radically more efficient drug development.Advancements in AI and drug discovery are creating more candidate drugs than the ind...Show moreLast updated: 3 days ago
    • Promoted
    Senior Data Engineer

    Senior Data Engineer

    Garner HealthNew York, NY, United States
    Full-time
    Healthcare quality is declining and soaring costs are crushing American families and businesses.At Garner, we've developed a revolutionary approach to evaluating doctor performance and a unique inc...Show moreLast updated: 30+ days ago
    • Promoted
    Data Engineer III (US)

    Data Engineer III (US)

    TD BankNew York, NY, United States
    Full-time
    New York, New York, United States of America.TD is committed to providing fair and equitable compensation opportunities to all colleagues. Growth opportunities and skill development are defining fea...Show moreLast updated: 4 days ago
    • Promoted
    Senior Data Engineer

    Senior Data Engineer

    Publicis Groupe Holdings B.VNew York, NY, United States
    Full-time
    Zenith is a full-service media agency with capabilities and expertise across all channels and disciplines.Zenith is part of Publicis Media, the #1 media buying network in the Americas and #2 global...Show moreLast updated: 4 days ago
    • Promoted
    Data Engineer III

    Data Engineer III

    AmazonNew York, NY, United States
    Full-time
    Offered Position : Data Engineer III.Job Location : New York, New York.Design, develop, implement, test, document, and operate large-scale, high-volume, high-performance data structures for business ...Show moreLast updated: 4 days ago
    • Promoted
    Senior Data Engineer

    Senior Data Engineer

    ZipRecruiterNew York, NY, United States
    Full-time
    Job DescriptionJob Description .Maven is the world's largest virtual clinic for women and families on a mission to make healthcare work for all of us. Maven's award-winning digital programs provide ...Show moreLast updated: 4 days ago
    • Promoted
    Senior Data Engineer I (Databricks)

    Senior Data Engineer I (Databricks)

    ZipRecruiterNew York, NY, United States
    Part-time
    Job DescriptionJob DescriptionCompany Description.Curinos empowers financial institutions to make better, faster and more profitable decisions through industry-leading proprietary data, technologie...Show moreLast updated: 4 days ago
    • Promoted
    Senior Data Engineer

    Senior Data Engineer

    GoldbelyNew York, NY, United States
    Full-time
    At Goldbelly, we believe food brings people together.We connect people with their greatest culinary desires within and beyond local communities. We empower food makers of all sizes and deliver their...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Data Engineer

    Senior Data Engineer

    Publicis GroupeNew York, NY, United States
    Full-time
    Zenith is a full-service media agency with capabilities and expertise across all channels and disciplines.Zenith is part of Publicis Media, the #1 media buying network in the Americas and #2 global...Show moreLast updated: 4 days ago
    • Promoted
    Data Engineer III

    Data Engineer III

    The Custom Group of CompaniesNew York, NY, United States
    Full-time
    Your role as a Senior Data Engineer.Work on migrating applications from an on-premises location to the cloud service providers. Develop products and services on the latest technologies through contr...Show moreLast updated: 4 days ago
    • Promoted
    Data Engineer II

    Data Engineer II

    Horizon MediaNew York, NY, United States
    Full-time
    Bill Koenigsberg, is recognized as one of the most innovative marketing and advertising firms.We are headquartered in New York City, with offices in Los Angeles and Toronto.A leader in driving busi...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Data Engineer II (Databricks)

    Senior Data Engineer II (Databricks)

    ZipRecruiterNew York, NY, United States
    Part-time
    Job DescriptionJob DescriptionCompany Description.Curinos empowers financial institutions to make better, faster and more profitable decisions through industry-leading proprietary data, technologie...Show moreLast updated: 4 days ago
    • Promoted
    Data Engineer II

    Data Engineer II

    VirtualVocationsElizabeth, New Jersey, United States
    Full-time
    A company is looking for a Data Engineer II.Key Responsibilities Produce high-quality data models and maintain data integrity for analytics products Develop scalable ELT pipelines and business i...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Data Engineer

    Senior Data Engineer

    Starcom Mediavest Group Germany GmbhNew York, NY, United States
    Full-time
    Zenith is a full-service media agency with capabilities and expertise across all channels and disciplines.Zenith is part of Publicis Media, the #1 media buying network in the Americas and #2 global...Show moreLast updated: 4 days ago
    • Promoted
    Data Analytics Engineer II

    Data Analytics Engineer II

    Garner HealthNew York, NY, US
    Full-time
    Healthcare quality is declining and soaring costs are crushing American families and businesses.At Garner, we've developed a revolutionary approach to evaluating doctor performance and a unique...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Data Engineer

    Senior Data Engineer

    oritaNew York, NY, United States
    Full-time
    Direct-to-consumer brands pay us, in order to market less.Well, technically, they pay us to market more effectively.And, strangely (!), that often means marketing a lot less.How? We use a lot of ma...Show moreLast updated: 4 days ago