1

Duckdb Jobs in Virginia (NOW HIRING)

Jr. Campaign Analyst

Arlington, VA · On-site

$100 - $120/hr

Experience working with Python, R, MySQL/DuckDB, C++ * Experience/knowledge of Space Operations and their contribution to the joint warfight Please note that the salary information shown below is a ...

Be Seen First

Familiarity with Python packages like Pandas, Spark.sql, and DuckDB * Knowledge in DBT and understanding of DAGs * Hands-on experience with cloud platforms AWS, GCP, or Azure * Experience in data ...

Duckdb information

What is DuckDB?

DuckDB is an in-process SQL OLAP (Online Analytical Processing) database management system designed for fast analytical data processing. Unlike traditional client-server databases, DuckDB runs directly within your application and works well with data science tools and workflows. It is lightweight, easy to install, and supports SQL queries on data stored in CSV, Parquet, and other formats, making it popular for data analysis and research use cases.

What are some common challenges faced by professionals working with DuckDB in data engineering roles?

Professionals using DuckDB in data engineering often encounter challenges such as optimizing query performance for large-scale datasets, integrating DuckDB with existing data pipelines, and ensuring compatibility with various data formats and sources. Additionally, since DuckDB is relatively new compared to other database systems, users may find limited community support or resources for advanced use cases. Collaborating with data scientists and analysts is key, as DuckDB is often used for interactive analytics, requiring close communication to tailor solutions that meet both performance and usability needs.

What are the key skills and qualifications needed to thrive as a DuckDB developer, and why are they important?

To thrive as a DuckDB Developer, you need strong SQL proficiency, a background in database management, and experience with analytical data processing, ideally supported by a degree in computer science or a related field. Familiarity with DuckDB, data integration tools, and programming languages such as Python or R is essential. Attention to detail, problem-solving abilities, and effective communication are valuable soft skills in this role. These skills ensure efficient data analysis, seamless database integration, and the ability to convey insights to technical and non-technical stakeholders.

What is the difference between Duckdb vs Data Analyst?

AspectDuckdbData Analyst
Primary RoleEmbedded analytical database engine for data processingInterprets data, creates reports, and provides insights
Required SkillsSQL, data management, database integrationData visualization, statistical analysis, SQL
Work EnvironmentDevelopers, data engineers, embedded systemsBusiness environments, analytics teams, offices
CertificationsNone specific, SQL knowledge preferredCertified Data Analyst, SQL certifications

Duckdb is a database engine used mainly by developers and data engineers for embedded data processing, while Data Analysts focus on interpreting data and generating insights. Both roles require SQL skills, but their work environments and objectives differ significantly.

What can I do with DuckDB?

As a data analyst or developer, you can use DuckDB to perform efficient in-process SQL queries on large datasets, often within Python or R environments. It supports complex analytical queries, data transformation, and integration with data science workflows, making it suitable for data analysis, machine learning, and reporting tasks.

What cities in Virginia are hiring for Duckdb jobs?

Cities in Virginia with the most Duckdb job openings:

Infographic showing various Duckdb job openings in Virginia as of August 2026, with employment types broken down into 98% Full Time, and 2% Contract. Highlights an 82% Physical, 5% Hybrid, and 13% Remote job distribution.

Senior Backend Engineer (Agentic Data Platform)

Namely

Portsmouth, VA • On-site

$140 - $190/hr

Other

Posted 19 days ago


Key responsibilities

  • Build and improve distributed backend systems for the genomics platform.

  • Develop and optimize data processing pipelines over large genomic and health datasets using Apache Spark and DuckDB.

  • Design data structures, storage, and indexing strategies across PostgreSQL, Qdrant, and Redis to ensure performance at scale.


Job description

The opportunity

As a Senior Backend Engineer, you'll build and improve the distributed backend systems behind our genomics platform. At its core, this platform is a high-throughput data processing and retrieval system with an AI-powered natural language interface - you'll work with our AI Backend Architect on the services and pipelines that make genetics-based guidance fast, accurate, and reliable. When you do this well, people can have meaningful conversations with their DNA and receive trustworthy guidance that evolves alongside advances in science and AI.

This is a backend and data engineering role. It is not an LLM-integration role, and it is not about adding AI tooling to a product.

What you'll own
  • Develop and maintain high-performance and fault-tolerant distributed backend services.
  • Build and optimize data processing pipelines over large genomic and health datasets using Apache Spark and DuckDB.
  • Design data structures, storage and indexing strategies across PostgreSQL, Qdrant, and Redis for performance at scale.
  • Own async event processing and workflow orchestration using AWS SQS, RabbitMQ, and BullMQ.
  • Drive latency, throughput, and reliability through parallel execution, caching, and efficient data access.
  • Build guardrail and quality assurance layers that keep AI responses anchored in real genomic and scientific evidence.
  • Partner with bioinformatics experts to ensure outputs match the science, and with product/design specialists on user-facing behavior.
Who you are
  • 5+ years building and operating production backend systems at scale.
  • Expert-level in TypeScript, comfortable owning production services end to end.
  • Strong distributed-systems fundamentals - you understand how they're designed and why they fail.
  • Hands-on with large-scale data processing frameworks (Apache Spark or equivalent) and very large datasets.
  • Deeply familiar with both OLTP and OLAP data systems (PostgreSQL, DuckDB)
  • Solid with distributed event-driven systems.
  • Able to step into an unfamiliar domain like genomics, learn the mechanics fast, and go deep.
  • Craft-driven: you build systems properly with attention to detail and high bar for quality rather than assembling pre-made pieces.
  • This is a fully remote role open to candidates in time zones from UTC−5 to UTC+3.
Bonus to have
  • Rust, and functional programming experience (Scala or similar).
  • Python for data processing.
  • Experience training or fine-tuning your own AI models - not just calling APIs.
  • Experience with multi-agent AI systems and orchestration (planner / router / evaluator patterns).
  • Production experience with LLM APIs (Anthropic, OpenAI, Bedrock, Google AI)
  • Production experience with vector databases and RAG pipelines.
  • LLM observability tooling (Langfuse, LangSmith).
  • Workflow engines (Temporal).
  • Familiarity with genomics, bioinformatics, or health data systems.
  • High-growth startup experience.
Why this role matters

AI is transforming how people access information. Genomics is transforming how people understand themselves.

This role sits at the convergence of both.

You’ll help build the AI systems that enable people to interact with their DNA and receive personalized guidance throughout their lives. The systems you build will help millions of people better understand their health, identify risks earlier, make more informed decisions, and benefit from advances in science that would otherwise remain inaccessible.

This is an opportunity to help create a category-defining, generational product and shape how humanity interacts with its DNA for decades to come.

#J-18808-Ljbffr