About the Team
RA Capital's Data Engineering team builds the enterprise data platform that powers research, investment, and operational workflows across the firm. Beyond traditional data pipelines, we are actively developing AI-native data access patterns that enable users to interact with complex healthcare datasets through LLM-powered, governed, and auditable interfaces.
Our work sits at the intersection of healthcare data, large-scale data engineering, and applied AI, with a strong emphasis on reliability, compliance, and production readiness.
About the Role
We are seeking an engineer to help build and operate AI-enabled data platforms for healthcare data.
This role goes beyond traditional ETL. You will work on healthcare data pipelines and contribute to AI-driven data access layers where large language models translate user intent into structured, governed queries, such as SQL or GraphQL. You will help bridge enterprise data systems with LLM-based interfaces that support self-service analytics, discovery, and decision-making.
Responsibilities
Design and maintain scalable, production-grade data pipelines for healthcare datasets, including claims and provider data.
Build and optimize data models in Databricks to support analytical and AI-driven workloads.
Develop high-quality Python and SQL code for data standardization, reconciliation, and validation.
Partner with AI/ML engineers to enable LLM-powered data access.
Contribute to AI-native data access frameworks.
Collaborate with healthcare data vendors and internal stakeholders.
Implement monitoring, validation, and governance for data and AI workflows.
Document data architectures and AI-enabled pipelines.
Key Skills and Experience
Strong foundation in data engineering and data modeling.
Proficiency in SQL, Python, Java, and Spark.
Hands-on experience with Databricks and/or Snowflake.
Experience in software development, data integration, vendor data management, and writing production-grade code for data reconciliation.
Experience working with healthcare or other regulated datasets.
Hands-on experience with, or a strong interest in, LLM-powered enterprise systems.
Familiarity with GraphQL or other API-based data access methods.
Strong documentation and communication skills.
Key Requirements
Bachelor's degree or higher in Computer Science, Data Science, Information Technology, Software Engineering, or a related field. A master's degree is preferred.
1-3+ years of experience in software or data engineering.
Authorization to work in the United States without current or future sponsorship.
Must be based in Massachusetts, able to work in the Boston area, and available to work a hybrid schedule. Relocation assistance is not provided.
Flexibility to work outside standard business hours as needed.