1

Hierarchical Clustering Jobs (NOW HIRING)

Senior AI/ML Engineer

Atlanta, GA · On-site

$100K - $138K/yr

Develop clustering algorithms (DBSCAN, hierarchical clustering) to create unified "golden customer profiles" that serve as the authoritative representation of each individual * Build embedding-based ...

AM Quantitative Analyst I

Boston, MA · On-site

$135K - $175K/yr

... hierarchical clustering and Centroid based - K-means algorithm), Bayesian statistics, Time-Series analysis, and non-linear tree-based models. * DE streamlining data preparation pipeline using ...

New

Data Scientist

Vancouver, WA · Hybrid

$65 - $70/hr

Unsupervised Learning (e.g., clustering techniques, hierarchical clustering, dimensionality reduction, principal component analysis). Time Series Analysis. Demonstrated knowledge of computer ...

Sr. Data Analyst

Orlando, FL

$80K - $101K/yr

Design, develop, and implement advanced algorithms and automated analytical pipelines, including clustering (e.g k-means, hierarchical clustering, DBSCAN, HDBSCAN), supervised learning (e.g. linear ...

Data Scientist

Vancouver, WA · On-site

$65 - $70/hr

Unsupervised Learning (e.g., clustering techniques, hierarchical clustering, dimensionality reduction, principal component analysis). Time Series Analysis. Demonstrated knowledge of computer ...

next page

Showing results 1-20

Hierarchical Clustering information

See salary details

$11K

$110K

$184.5K

How much do hierarchical clustering jobs pay per year?

As of Jul 20, 2026, the average yearly pay for hierarchical clustering in the United States is $109,999.00, according to ZipRecruiter salary data. Most workers in this role earn between $60,000.00 and $160,000.00 per year, depending on experience, location, and employer.

What is the difference between Hierarchical Clustering vs Data Analyst?

AspectHierarchical ClusteringData Analyst
Primary RoleUnsupervised machine learning technique for data groupingAnalyzing data to identify trends and support decision-making
Required SkillsStatistical analysis, programming (Python/R), understanding of clustering algorithmsData visualization, statistical analysis, Excel, SQL
Work EnvironmentData science teams, research projects, machine learning applicationsBusiness environments, reporting, data interpretation

Hierarchical Clustering is a machine learning method used to group similar data points without labeled outcomes, often employed in data science projects. Data Analysts focus on interpreting data, creating reports, and providing insights to inform business decisions. While both roles work with data, Hierarchical Clustering involves technical algorithm development, whereas Data Analysts focus on data interpretation and communication.

What are the applications of hierarchical clustering?

Hierarchical clustering is used in various fields such as bioinformatics for gene expression analysis, market research for customer segmentation, and image analysis for object recognition. It helps identify natural groupings in data without pre-specifying the number of clusters, making it useful for exploratory data analysis and pattern discovery. Skills in data preprocessing and familiarity with clustering tools like R or Python are beneficial for applying this method effectively.

What is an example of a job hierarchy?

A job hierarchy is a structured arrangement of roles within an organization, such as entry-level staff, team leaders, managers, directors, and executives. Hierarchies help define reporting relationships, responsibilities, and authority levels, often visualized in organizational charts. Hierarchical clustering in data analysis is unrelated to organizational structures but shares the concept of grouping similar items.

What are the 4 types of clustering?

In hierarchical clustering, the four main types are agglomerative, divisive, single linkage, and complete linkage. Agglomerative starts with individual data points and merges them, while divisive begins with the entire dataset and splits it into clusters. Single linkage merges clusters based on the closest points, and complete linkage considers the furthest points between clusters, affecting the shape and size of the resulting clusters.

What is a real life example of hierarchical clustering?

Hierarchical clustering is used in customer segmentation to group consumers based on purchasing behavior, enabling targeted marketing strategies. It can also be applied in document organization, such as organizing news articles or research papers into related topics, often using data analysis tools like R or Python. These applications help analysts identify natural groupings within complex data sets.
More about Hierarchical Clustering jobs
What states have the most Hierarchical Clustering jobs? States with the most job openings for Hierarchical Clustering jobs include:
What job categories do people searching Hierarchical Clustering jobs look for? The top searched job categories for Hierarchical Clustering jobs are:
Infographic showing various Hierarchical Clustering job openings in the United States as of July 2026, with employment types broken down into 86% Full Time, 9% Part Time, and 5% Contract. Highlights an 76% Physical, 5% Hybrid, and 19% Remote job distribution, with an average salary of $109,999 per year, or $52.9 per hour.
Senior AI/ML Engineer

Senior AI/ML Engineer

Sumeru

Atlanta, GA • On-site

$100K - $138K/yr

Other

Re-posted 4 days ago


Job description

Role: Senior AI/ML Engineer
Location: Bellevue/Seattle, WA ; Atlanta, GA, and Frisco, TX


Need Local Candidates


Job Overview

We are seeking an AI/ML Engineer to build the intelligent systems that power identity resolution and data accessibility within our Customer Data Platform (CDP) - the authoritative source of truth for customer data across the entire US adult population.

This role focuses on developing machine learning pipelines that deduplicate, link, and resolve customer identities across disparate data sources - the core capability that transforms raw data into trusted, unified customer profiles. You will also contribute to LLM-based solutions that enable natural language querying of CDP data, making the platform accessible to business users across the organization.

You will work on both classical ML techniques and modern LLM-based approaches to ensure that every customer identity in CDP is accurately resolved, every profile is trustworthy, and every user can access the data they need.

Job Responsibilities - Identity Resolution

  • Develop and deploy entity resolution models to match and deduplicate customer records across multiple systems - directly impacting the accuracy of CDP as the source of truth
  • Implement probabilistic matching techniques (e.g., Fellegi-Sunter) and ML models (gradient boosting, neural classifiers) for record linkage across the US adult population
  • Build candidate blocking pipelines using phonetic algorithms (Soundex, Double Metaphone), token similarity, and LSH to handle billions of potential match pairs efficiently
  • Apply fuzzy matching techniques (Levenshtein, Jaro-Winkler, Jaccard) for customer attributes such as name, address, phone, and identifiers
  • Develop clustering algorithms (DBSCAN, hierarchical clustering) to create unified "golden customer profiles" that serve as the authoritative representation of each individual
  • Build embedding-based similarity systems using Sentence-BERT or transformer-based models for semantic matching
  • Implement ANN/KNN retrieval systems (FAISS, Annoy) for large-scale entity matching across population-scale datasets

Job Responsibilities - AI/LLM

  • Use LLMs (e.g., GPT, Claude) for classification and disambiguation of entity matches, improving resolution accuracy where traditional methods fall short
  • Build and support RAG pipelines to enrich customer profiles with contextual data from unstructured sources
  • Perform prompt engineering and evaluation for structured data extraction from unstructured inputs feeding into CDP
  • Contribute to NLQ-to-SQL systems, enabling business users to query CDP data using natural language - making the authoritative source of truth accessible to non-technical stakeholders
  • Support integration with vector databases (e.g., Pinecone, pgvector, Qdrant) for semantic search across customer data

Education and Work Experience

  • Bachelor's or Master's degree in Computer Science, Data Science, or related field
  • 3+ years of experience in ML/AI engineering
  • At least 1 year of experience in entity resolution, record linkage, or deduplication - ideally at scale

Technical Skills

  • Programming: Python (required)
  • Libraries: scikit-learn, HuggingFace Transformers, RapidFuzz, jellyfish
  • Experience with LLM APIs (OpenAI, Anthropic) and prompt pipelines
  • Strong SQL skills and experience with Spark or Dask for distributed processing
  • Familiarity with vector databases and embedding-based retrieval
  • Experience with ML lifecycle tools (MLflow or similar)
  • Understanding of data quality metrics and how identity resolution impacts downstream trust

Knowledge, Skills, and Abilities

  • Strong understanding of ML fundamentals and similarity matching techniques applied to customer identity
  • Ability to work with large, messy, real-world datasets spanning hundreds of millions of records
  • Understanding of precision/recall tradeoffs in identity resolution and their impact on data trust
  • Good problem-solving and analytical skills
  • Ability to collaborate with data engineering, platform, and business teams to deliver accurate customer profiles

Sumeru logo

About Sumeru

Sourced by ZipRecruiter

Industry

It services

Company size

501 - 1,000 Employees

Headquarters location

Washington, DC, US

Year founded

2002