1

Linked Data Jobs (NOW HIRING)

Data Engineer (Founding Team)

Bodega Bay, CA · On-site

$135K - $163K/yr

Palantir Ontology, Linked Data, W3C standards) * Familiar with fine-tuning LLMs or enabling RAG pipelines using enterprise knowledge * Experience enforcing data access policy with tools like OPA ...

Associate Data/ Metadata Analyst

Dublin, OH · On-site

  • Medical

  • Retirement

Emerging Tech: Interest in applying AI and linked data concepts (e.g., Schema.org) to library data challenges. Working Conditions: Normal office environment. ADA/EAA: The above statements cover what ...

Principal Data Scientist - Oncology

Titusville, NJ

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Programming background in parser combinators, natural language processing, and linked data (RDF Triple Stores and property graphs). * Proficiency in semantic web technologies (e.g. SPARQL, RDF, OWL ...

Principal Data Scientist - Oncology

San Diego, CA

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Programming background in parser combinators, natural language processing, and linked data (RDF Triple Stores and property graphs). * Proficiency in semantic web technologies (e.g. SPARQL, RDF, OWL ...

Principal Data Scientist - Oncology

Spring House, PA

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Programming background in parser combinators, natural language processing, and linked data (RDF Triple Stores and property graphs). * Proficiency in semantic web technologies (e.g. SPARQL, RDF, OWL ...

Principal Data Scientist - Oncology

Raritan, NJ

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Programming background in parser combinators, natural language processing, and linked data (RDF Triple Stores and property graphs). * Proficiency in semantic web technologies (e.g. SPARQL, RDF, OWL ...

Principal Data Scientist - Oncology

Cambridge, MA

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Programming background in parser combinators, natural language processing, and linked data (RDF Triple Stores and property graphs). * Proficiency in semantic web technologies (e.g. SPARQL, RDF, OWL ...

Showing results 21-40

Linked Data information

What is linked data?

Linked Data refers to a method of publishing structured data on the web so that it can be easily connected and queried across different sources. It uses standard web technologies such as HTTP, RDF, and URIs to enable data from different domains to be linked and integrated. This approach allows for greater interoperability and discoverability of information, making it easier to build applications that use data from diverse sources. Linked Data plays a key role in the Semantic Web, supporting more intelligent and context-aware web services.

What are some common challenges faced when working as a linked data specialist, and how can they be addressed?

Linked Data specialists often encounter challenges such as integrating heterogeneous data sources, ensuring data quality, and maintaining semantic consistency across datasets. Addressing these issues typically involves using established ontologies, adhering to best practices in RDF modeling, and collaborating closely with domain experts and data owners. Regular team meetings and documentation help ensure consistency, while leveraging open standards and validation tools can minimize errors and incompatibilities.

What are the key skills and qualifications needed to thrive as a linked data specialist, and why are they important?

To thrive as a Linked Data Specialist, you need a solid background in data modeling, semantic web technologies, and knowledge of standards like RDF and SPARQL, often supported by a degree in computer science or information science. Familiarity with tools such as Protégé, triple stores (e.g., Apache Jena, Virtuoso), and ontology editors, as well as experience with web data integration, is typically required. Strong analytical thinking, attention to detail, and effective communication skills help distinguish top performers in this role. These skills enable accurate data linking, interoperability, and the development of robust, reusable semantic data solutions across organizations.

What is the difference between Linked Data vs Data Analyst?

AspectLinked DataData Analyst
Required CredentialsKnowledge of RDF, SPARQL, ontologiesBachelor's in statistics, data science, or related field
Work EnvironmentSemantic web projects, data integration, knowledge graphsData interpretation, reporting, business insights
Employer & Industry UsageTech, research, semantic web companiesFinance, marketing, healthcare, various industries
Search & Comparison IntentUnderstanding semantic data structuresAnalyzing and interpreting data sets

While both roles involve working with data, Linked Data focuses on structuring and connecting data using semantic web technologies, whereas Data Analysts interpret data to provide business insights. The roles differ in tools, environment, and objectives but share a common goal of leveraging data effectively.

More about Linked Data jobs
Infographic showing various Linked Data job openings in the United States as of August 2026, with employment types broken down into 1% As Needed, 84% Full Time, 11% Part Time, and 4% Contract. Highlights an 86% Physical, 4% Hybrid, and 10% Remote job distribution.

Data Engineer (Founding Team)

Fabrion

Bodega Bay, CA • On-site

$135K - $163K/yr

Full-time

Re-posted 9 days ago


Job description

Data/ETL Engineer (Founding Team)

Location: San Francisco Bay Area

Type: Full-Time

Compensation: Competitive salary + early-stage equity

Backed by 8VC, we're building a world-class team to tackle one of the industry’s most critical infrastructure problems.

About the Role

We’re building a multi-tenant, AI-native platform where enterprise data becomes actionable through semantic enrichment, intelligent agents, and governed interoperability. At the heart of this architecture lies our Data Fabric — an intelligent, governed layer that turns fragmented and siloed data into a connected ontology ready for model training, vector search, and insight-to-action workflows.

We're looking for engineers who enjoy hard data problems at scale: messy unstructured data, schema drift, multi-source joins, security models, and AI-ready semantic enrichment. You’ll build the backend systems, data pipelines, connector frameworks, and graph-based knowledge models that fuel agentic applications.

If you've worked on streaming unstructured pipelines, built connectors into ugly legacy systems, or mapped knowledge graphs that scale — this role will feel like home.

Responsibilities
  • Build highly reliable, scalable data ingestion and transformation pipelines across structured, semi-structured, and unstructured data sources

  • Develop and maintain a connector framework for ingesting from enterprise systems (ERPs, PLMs, CRMs, legacy data stores, email, Excel, docs, etc.)

  • Design and maintain the data fabric layer — including a knowledge graph (Neo4j or Puppygraph) enriched with ontologies, metadata, and relationships

  • Normalize and vectorize data for downstream AI/LLM workflows — enabling retrieval-augmented generation (RAG), summarization, and alerting

  • Create and manage data contracts, access layers, lineage, and governance mechanisms

  • Build and expose secure APIs for downstream services, agents, and users to query enriched semantic data

  • Collaborate with ML/LLM teams to feed high-quality enterprise data into model training and tuning pipelines

What We’re Looking For

Core Experience:

  • 5+ years building large-scale data infrastructure in production environments

  • Deep experience with ingestion frameworks (Kafka, Airbyte, Meltano, Fivetran) and data pipeline orchestration (Airflow, Dagster, Prefect)

  • Comfortable processing unstructured data formats: PDFs, Excel, emails, logs, CSVs, web APIs

  • Experience working with columnar stores, object storage, and lakehouse formats (Iceberg, Delta, Parquet)

  • Strong background in knowledge graphs or semantic modeling (e.g. Neo4j, RDF, Gremlin, Puppygraph)

  • Familiarity with GraphQL, RESTful APIs, and designing developer-friendly data access layers

  • Experience implementing data governance: RBAC, ABAC, data contracts, lineage, data quality checks

Mindset & Culture Fit:

  • You’re a system thinker: you want to model the real world, not just process it

  • Comfortable navigating ambiguous data models and building from scratch

  • Passionate about enabling AI systems with real-world, messy enterprise data

  • Pragmatic about scalability, observability, and schema evolution

  • Value autonomy, high trust, and meaningful ownership over infrastructure

Bonus Skills

  • Prior work with vector DBs (e.g. Weaviate, Qdrant, Pinecone) and embedding pipelines

  • Experience building or contributing to enterprise connector ecosystems

  • Knowledge of ontology versioning, graph diffing, or semantic schema alignment

  • Familiarity with data fabric patterns (e.g. Palantir Ontology, Linked Data, W3C standards)

  • Familiar with fine-tuning LLMs or enabling RAG pipelines using enterprise knowledge

  • Experience enforcing data access policy with tools like OPA, Keycloak, Snowflake row-level security

Why This Role Matters

Agents are only as smart as the data they operate on. This role builds the foundation — the semantic, governed, connected substrate — that makes autonomous decision-making and agent action possible. From factory ERP records to geopolitical news alerts, the data fabric unifies it all.

If you're excited to tame complexity, unify chaos, and power intelligent systems with trusted data — we’d love to hear from you.