1

Capacity Engineer Jobs (NOW HIRING)

Meta is seeking a Data Center Capacity Engineer to support the planning, analysis, and optimization of server and infrastructure capacity across our global data center fleet. In this role, you will ...

BIM Engineer

Austin, TX · On-site

$164K - $206K/yr

Capacity Delivery owns technical execution across every build: the drawings, the field questions, and the engineering calls that keep multiple concurrent gigawatt sites moving. * Engineer data ...

Supply Chain Capacity Engineer Responsibilities: * Own infrastructure supply chain planning and operations for Meta (as part of Meta's overall capacity plan): including servers, network, and data ...

Supply Chain Capacity Engineer Responsibilities: * Own infrastructure supply chain planning and operations for Meta (as part of Meta's overall capacity plan): including servers, network, and data ...

Meta is seeking a Supply Chain Engineer to join the Infrastructure Supply Chain and Engineering ... balance capacity, flexibility, supply capability, and cost, from base components through ...

Supply Chain Capacity Engineer Responsibilities: * Own infrastructure supply chain planning and operations for Meta (as part of Meta's overall capacity plan): including servers, network, and data ...

GM is looking for aSeniorCapacity Engineer to join the AV Capacity and Performance Engineering teamin the AVInfrastructure orgto support our critical efforts in developing autonomous vehicles.The ...

Showing results 21-40

Capacity Engineer information

See salary details

$53K

$127K

$143.5K

How much do capacity engineer jobs pay per year?

As of Sep 1, 2026, the average yearly pay for capacity engineer in the United States is $127,019.00, according to ZipRecruiter salary data. Most workers in this role earn between $116,000.00 and $143,000.00 per year, depending on experience, location, and employer.

What is a capacity engineer?

A Capacity Engineer is responsible for analyzing, planning, and optimizing system resources to ensure efficient performance and scalability. They monitor infrastructure utilization, forecast future capacity needs, and collaborate with teams to prevent resource shortages or over-provisioning. Capacity Engineers work with hardware, software, and cloud environments to maintain system reliability while balancing cost-effectiveness. They use data analysis, performance metrics, and automation tools to ensure optimal system efficiency.

What does a capacity engineer do?

As a Capacity Engineer, your daily responsibilities often include monitoring system performance, analyzing usage data, and forecasting future resource needs to ensure infrastructure remains robust and efficient. You'll regularly collaborate with IT, network, and business teams to align capacity planning with organizational objectives and upcoming projects. The role also involves preparing reports, recommending upgrades or adjustments, and troubleshooting any emerging capacity issues. This combination of technical analysis and cross-functional teamwork ensures smooth operations and supports the organization's scalability.

What are the key skills and qualifications needed to thrive as a capacity engineer?

To thrive as a Capacity Engineer, you need a solid background in systems engineering, network design, and data analytics, often supported by a degree in engineering or a related field. Familiarity with capacity planning tools, performance monitoring platforms, and relevant certifications such as ITIL or CCNA are commonly required. Strong problem-solving skills, attention to detail, and effective communication abilities help set top candidates apart. These skills ensure accurate forecasting, resource optimization, and clear collaboration across technical and business teams.

More about Capacity Engineer jobs

What are the most commonly searched types of Capacity Engineer jobs?

The most popular types of Capacity Engineer jobs are:

Infographic showing various Capacity Engineer job openings in the United States as of August 2026, with employment types broken down into 93% Full Time, 3% Part Time, and 4% Contract. Highlights an 86% Physical, 5% Hybrid, and 9% Remote job distribution, with an average salary of $127,019 per year, or $61.1 per hour.

Staff+ Software Engineer, Capacity Engineering

San Francisco, CA • On-site

Anthropic
Software Development • 11 - 50 employees

Full-time

Re-posted 18 days ago


Job description

About the Role

Anthropic manages one of the largest and fastest-growing infrastructure fleets in the industry - spanning multiple accelerator families, cpu families and clouds. The Capacity Engineering team is responsible for making sure all our infrastructure resources are accounted for, well-utilized, and efficiently allocated. We own the data, tooling, and operational systems that let Anthropic plan, measure, and maximize utilization across first-party and third-party compute.

As an engineer on Capacity Engineering, you will build the production systems that power this work: data pipelines that ingest and normalize telemetry from heterogeneous cloud environments, observability tooling that gives the org real-time visibility into fleet health, and performance instrumentation that measures how efficiently every major workload uses the hardware it's running on. You will be expected to write production-quality code every day, operate alongside Kubernetes-native infrastructure at meaningful scale, and directly influence decisions around one of Anthropic's largest areas of spend.

You'll collaborate closely with research engineering, infrastructure, inference, and finance teams. The work requires someone who can move between data engineering, systems engineering, and observability with comfort - and who thrives in a high-autonomy, high-ambiguity environment.

This is a pipeline role feeding four areas. Depending on your background and business priority, you'll focus primarily in one, but the boundaries are fluid and the problems overlap:

  • Data platform Pipelines that ingest occupancy and utilization telemetry from Kubernetes clusters, normalize billing and usage across cloud providers, and serve the BigQuery tables the rest of the org queries against. Correctness, completeness, and latency are the job, not a footnote. Consumers range from research engineers to finance to leadership, so it's product work as much as engineering: defining schema contracts, making data discoverable, and figuring out what people actually need.
  • Planning Knowing what the fleet has, where it's going, and what's in the way. Making the state of the fleet legible and actionable in real time: cluster health tooling, capacity planning platforms, alerting on occupancy drops and allocation problems, and systemic fixes to scheduling and fragmentation. Kubernetes operations on one side, cross-team coordination on the other.
  • Efficiency Measuring and improving how effectively every major workload uses the hardware it runs on. Instrumenting utilization across training, inference, and eval systems, building benchmarking infrastructure, establishing per-config baselines, and working directly with system-owning teams to close the gaps. The metric has to be good enough that the team on the hook for it agrees with the number.
  • Attribution and forecasting Connecting what the fleet costs to what the business is doing with it. Reconciling CSP billing exports against vendor telemetry and internal systems with mismatched schemas, attributing spend to the workloads and teams that generate it, and turning inference demand signals and research roadmaps into a defensible compute plan. Efficiency metrics have to survive contact with finance: stripped of pure demand and unit-price effects, reproducible month over month, and legible to a CFO.
Key responsibilities
  • Build the planning and allocation stack - the tools leadership uses to allocate capacity, teams use to plan against their allocations, and the scheduler enforces. Cross-region and cross-provider placement, guardrails, queueing, occupancy KPIs.
  • Drive the efficiency programs: stranding and rightsizing, unused capacity recovery, and job-level utilization across training, inference, and eval. Establish per-config baselines and work with system-owning teams to close the gaps. Utilization improvements are worth enormous sums at our scale.
  • Own attribution and forecasting - reconcile billing across ten-plus providers against telemetry and internal systems, attribute spend to the workloads that generate it, and turn demand signals and research roadmaps into a defensible compute plan and supply pipeline.
  • Build the data platform underneath all of it: pipelines ingesting occupancy, utilization, and cost from a rapidly diversifying fleet into BigQuery, with real ownership of completeness, latency SLOs, and gap detection. Every new provider is a net-new integration.
  • Operate Kubernetes-native systems at scale - collection agents, workload labeling, and the taint/reservation/scheduling behavior that determines what capacity is actually usable.
  • Treat the output as a product, not a pipeline. Gather your own requirements, define schema contracts, and design for consumers ranging from research engineers to a CFO - including on-call and SLOs, because these surfaces are load-bearing for the company.
What you bring
  • A strong track record building and operating production systems. This is a hands-on engineering role with a devops flavor.
  • Python and SQL at production quality. Most pipeline code is Python; the presentation layer is BigQuery SQL, including table-valued functions and views. Both need to be idiomatic, well-tested, and maintainable.
  • Deep experience with at least one major cloud provider (Amazon Web Services, Google Cloud, or Microsoft Azure) and its operations
  • Experience with observability tooling stack, including Prometheus, PromQL, and Grafana, including writing recording rules and building monitoring that engineering teams rely on.
  • Ability to gather your own requirements and work across organizational boundaries in an ambiguous environment with limited direction.
Preferred qualifications
  • Experience with capacity planning, resource management, or cost attribution systems at a hyperscaler or in a large-scale machine learning environment. Time spent in product engineering and developer experience absolutely counts here.
  • Scheduling and packing efficiency experience, or profiling-driven optimization of large distributed workloads.
  • Multi-cloud data ingestion experience, especially normalizing billing exports, reservation APIs, on-demand capacity reservations, commitments, and vendor telemetry from providers with different billing arrangements.
  • Total cost of ownership and forecasting experience, including decomposing whether infrastructure growth is causal or correlated with business drivers.
  • Accelerator infrastructure familiarity. GPU metrics (DCGM), TPU utilization, Trainium power and utilization metrics, or experience with machine learning training and inference systems at the hardware level.
  • Experience building internal data products with self-service access, schema contracts, API serving, documentation, and discoverability. Not just pipelines, but thinking about how the data gets consumed.
  • Storage efficiency, retention, and lifecycle program experience at exabyte scale.