1

Duckdb Jobs in Massachusetts (NOW HIRING)

Sr Data Engineer

Boston, MA · Hybrid

$124K - $149K/yr

Advanced proficiency in SQL and Python (delta-rs, PyArrow, Polars, DuckDB, Pandas, and NumPy). * Proven experience designing modern data architectures (lakehouse, medallion, etc.). * Hands-on ...

Familiarity with modern data warehouses (e.g., Snowflake) and/or analytical engines (e.g., DuckDB/Polars) * Cloud experience, preferably Azure * Familiarity with containers and orchestration (e.g ...

Sr Data Engineer

Boston, MA · Hybrid

$124K - $149K/yr

Advanced proficiency in SQL and Python (delta-rs, PyArrow, Polars, DuckDB, Pandas, and NumPy). * Proven experience designing modern data architectures (lakehouse, medallion, etc.). * Hands-on ...

Modern data-stack tools (Airflow, Dagster, dbt, Snowflake/DuckDB/Iceberg) * Automated data testing (Great Expectations, dbt tests) * Experience with Notion, Nominal, or JMP * Additional DAQ platform ...

SQL fluency -- you think in queries, use DuckDB, dbt, or similar without looking things up; proficiency in Python preferred; comfortable reading and writing API integrations * Hands‑on experience ...

Data Engineer

Somerville, MA · On-site

$110 - $150/hr

We run Argo Workflows and Metaflow on EKS, Glue and Athena over Apache Iceberg and Parquet, DuckDB, and Terraform. Depth in any comparable stack transfers fine. * Demonstrated experience owning a ...

Data Engineer

Somerville, MA · On-site

$125K - $150K/yr

We run Argo Workflows and Metaflow on EKS, Glue and Athena over Apache Iceberg and Parquet, DuckDB, and Terraform. Depth in any comparable stack transfers fine. * Demonstrated experience owning a ...

Duckdb information

What is DuckDB?

DuckDB is an in-process SQL OLAP (Online Analytical Processing) database management system designed for fast analytical data processing. Unlike traditional client-server databases, DuckDB runs directly within your application and works well with data science tools and workflows. It is lightweight, easy to install, and supports SQL queries on data stored in CSV, Parquet, and other formats, making it popular for data analysis and research use cases.

What are some common challenges faced by professionals working with DuckDB in data engineering roles?

Professionals using DuckDB in data engineering often encounter challenges such as optimizing query performance for large-scale datasets, integrating DuckDB with existing data pipelines, and ensuring compatibility with various data formats and sources. Additionally, since DuckDB is relatively new compared to other database systems, users may find limited community support or resources for advanced use cases. Collaborating with data scientists and analysts is key, as DuckDB is often used for interactive analytics, requiring close communication to tailor solutions that meet both performance and usability needs.

What are the key skills and qualifications needed to thrive as a DuckDB developer, and why are they important?

To thrive as a DuckDB Developer, you need strong SQL proficiency, a background in database management, and experience with analytical data processing, ideally supported by a degree in computer science or a related field. Familiarity with DuckDB, data integration tools, and programming languages such as Python or R is essential. Attention to detail, problem-solving abilities, and effective communication are valuable soft skills in this role. These skills ensure efficient data analysis, seamless database integration, and the ability to convey insights to technical and non-technical stakeholders.

What is the difference between Duckdb vs Data Analyst?

AspectDuckdbData Analyst
Primary RoleEmbedded analytical database engine for data processingInterprets data, creates reports, and provides insights
Required SkillsSQL, data management, database integrationData visualization, statistical analysis, SQL
Work EnvironmentDevelopers, data engineers, embedded systemsBusiness environments, analytics teams, offices
CertificationsNone specific, SQL knowledge preferredCertified Data Analyst, SQL certifications

Duckdb is a database engine used mainly by developers and data engineers for embedded data processing, while Data Analysts focus on interpreting data and generating insights. Both roles require SQL skills, but their work environments and objectives differ significantly.

What can I do with DuckDB?

As a data analyst or developer, you can use DuckDB to perform efficient in-process SQL queries on large datasets, often within Python or R environments. It supports complex analytical queries, data transformation, and integration with data science workflows, making it suitable for data analysis, machine learning, and reporting tasks.

What cities in Massachusetts are hiring for Duckdb jobs?

Cities in Massachusetts with the most Duckdb job openings:

Infographic showing various Duckdb job openings in Massachusetts as of August 2026, with employment types broken down into 98% Full Time, and 2% Contract. Highlights an 82% Physical, 5% Hybrid, and 13% Remote job distribution.

$124K - $149K/yr

Contractor

This job post has expired today. Applications are no longer accepted.


Job description

The Data Engineer is responsible for building, and operating scalable, secure, and high-performance data platforms that enable analytics, reporting, and AI initiatives. This role leads the development of modern (ETL/ELT) data solutions across Microsoft Fabric and Azure, ensuring enterprise-grade data integration, governance, and performance optimization.

Top Must-Haves:

  • Azure Data Engineering Expertise - 5+ years building production data solutions on Azure, with strong hands-on experience in Microsoft Fabric plus tools like Azure Data Factory, Synapse, Databricks, or Azure SQL.
  • Strong ETL/ELT & Data Architecture Skills - Proven ability to design scalable pipelines and modern data architectures (lakehouse/medallion) supporting structured and unstructured data.
  • Advanced Programming & Data Processing - High proficiency in SQL and Python, plus experience with Spark and data processing frameworks (Pandas, PyArrow, etc.).
  • CI/CD & Data Platform Engineering - Experience implementing CI/CD pipelines, Git workflows, automated deployments, and environment promotion for data platforms.
  • Security & Enterprise Data Controls - Hands-on experience with RBAC, Key Vault, managed identities, and implementing secure, governed data architectures.

Preferred Skills:

  • Microsoft Fabric Certifications - DP-700 (required/expected), plus DP-600 or DP-203.
  • Performance Optimization & Cost Efficiency - Experience with Delta storage, partitioning, columnstore optimization, and query performance tuning.
  • Integration & Hybrid Data Environments - Experience integrating APIs, third-party platforms, and hybrid (on-prem + cloud) systems.
  • Analytics & AI Data Enablement - Experience supporting downstream semantic models and partnering with analytics/AI teams.
  • High-Scale / Regulated Environments - Background working in enterprise, highly regulated, or large-scale data platforms.
  • Leadership & Collaboration - Ability to lead design decisions, mentor junior engineers, and collaborate across technical and business teams in Agile environments.

Title: Sr. Data Engineer

Location: Woburn, MA - ONSITE 5 days a week

Duration: 6 month contract

Job Details:

  • Develop and maintain scalable ETL (Extract, Transform, Load) processes to efficiently extract data from diverse sources, transform it as required and load it into data warehouses or analytical systems.
  • Design and optimize database architectures and data pipelines to ensure high performance, availability and security while supporting structured and unstructured data.
  • Build and maintain robust ETL/ELT pipelines using Fabric Pipelines, Azure Data Factory, and Synapse.
  • Implement secure data architectures using Roles-based Access Controls (RBAC), Key Vault, Private Endpoints
  • Integrate data from APIs, third-party platforms, and hybrid (on-prem/cloud) systems.
  • Develop data workflows using Python, SQL, and Spark, selecting appropriate frameworks based on workload characteristics.
  • Drive cost-efficient, low-latency analytics through partition-aligned Delta storage, columnstore-optimized warehouse tables, DirectLake semantic access, and SCD-managed dimensional models, ensuring predicate pushdown, partition elimination, and minimal data movement across the query execution lifecycle.
  • Implement secure data architectures using:
    • RBAC and row/column/object-level security (RLS/CLS/OLS)
    • Azure Key Vault, Private Endpoints, Managed Identities.
  • Implement CI/CD pipelines for data platforms using Azure DevOps or GitHub
  • Establish automated deployment, environment promotion, and testing strategies.
  • Support downstream semantic models and reporting layers by delivering well-modeled, performant, and governed data structures (no report/dashboard development responsibilities).
  • Partner with analytics and AI teams to deliver trusted, production-ready data products.
  • Ensure platform reliability through constraint-driven data validation, fault-tolerant pipeline orchestration, and telemetry-backed observability, enabling anomaly detection, automated alerting, and lineage-driven root cause analysis across distributed data workloads.
  • Lead design decisions and influence enterprise data architecture standards
  • Collaborate with engineers, analysts, and business stakeholders to translate requirements into scalable solutions.
  • Mentor junior engineers and contribute to engineering excellence and knowledge sharing.
  • Operate effectively in Agile delivery environments.

Required Qualifications

  • 8+ years of experience in data, software, or platform engineering.
  • 5+ years building production data solutions on the Microsoft/Azure stack.
  • Strong experience with Microsoft Fabric and at least two of:
    • Azure Data Factory, Synapse, Databricks, Azure SQL, Power BI
  • Advanced proficiency in SQL and Python (delta-rs, PyArrow, Polars, DuckDB, Pandas, and NumPy).
  • Proven experience designing modern data architectures (lakehouse, medallion, etc.).
  • Hands-on experience with CI/CD, Git workflows, and environment promotion.
  • Experience implementing enterprise security, identity, and access controls.
  • Strong troubleshooting, performance tuning, and root cause analysis skills.
  • Experience working in regulated or high-scale environments.

Preferred Qualifications

  • Certifications:
    • DP-700 (Fabric Data Engineer) - required or within 6 months.
    • DP-600,DP-203- preferred.