1

Data Center Operations Engineer Jobs (NOW HIRING)

We're looking for a Staff Data Center Operations Engineer to serve as the senior technical operations resource for the SiteOps org -- based at our Denver headquarters, with cross-site scope and ...

... with data/reporting platforms (Cognos, Power BI) • Ability to analyze and troubleshoot complex, multi-layer issues end-to-end • Enterprise WMS systems, with Manhattan experience (Optional ...

Showing results 21-40

Data Center Operations Engineer information

See salary details

$51.5K

$147.5K

$197K

How much do data center operations engineer jobs pay per year?

As of Sep 6, 2026, the average yearly pay for data center operations engineer in the United States is $147,461.00, according to ZipRecruiter salary data. Most workers in this role earn between $84,000.00 and $196,000.00 per year, depending on experience, location, and employer.

What is a data center operations engineer?

A Data Center Operations Engineer is responsible for maintaining the daily operations of a data center, ensuring its infrastructure, servers, and networks run efficiently. They monitor system performance, troubleshoot hardware and software issues, and implement security and backup measures. Additionally, they collaborate with IT teams to optimize performance and support business continuity. Their role is crucial in minimizing downtime and ensuring data center reliability.

What are the typical daily responsibilities of a data center operations engineer?

As a Data Center Operations Engineer, your daily tasks often include monitoring system performance, performing preventative maintenance on servers and networking equipment, and responding quickly to hardware or connectivity issues. You'll be responsible for diagnosing and resolving operational incidents, managing backups, and applying security updates as needed. Collaboration with network, facilities, and IT teams is frequent to ensure maintenance activities and deployments are coordinated smoothly. This role can also involve creating documentation and managing ticket queues to keep operations running efficiently.

What are the key skills and qualifications needed to thrive as a data center operations engineer?

To thrive as a Data Center Operations Engineer, you need strong knowledge of server hardware, network infrastructure, troubleshooting, and systems administration, often supported by a degree in IT or a related field. Proficiency with monitoring tools (such as Nagios, SolarWinds), data center infrastructure management (DCIM) platforms, and certifications like CompTIA Server+, Cisco CCNA, or relevant vendor certifications are highly valued. Attention to detail, excellent problem-solving abilities, and effective communication skills help you excel when coordinating with cross-functional teams and responding to incidents. These skills ensure optimal data center performance, minimize downtime, and support the seamless delivery of IT services.

How much do data center operations engineers get paid?

Data center operations engineers typically earn between $60,000 and $100,000 annually, depending on experience, location, and certifications such as Cisco or CompTIA. Entry-level roles may start lower, while experienced engineers with specialized skills can earn higher salaries, often with opportunities for overtime and shift differentials.

What does a data center operations engineer do?

A data center operations engineer is responsible for maintaining and managing the infrastructure of data centers, including servers, networking equipment, power systems, and cooling. They monitor system performance, troubleshoot issues, perform routine maintenance, and ensure the security and reliability of data center operations, often using tools like monitoring software and following industry standards. Strong technical skills, problem-solving abilities, and certifications such as CompTIA Server+ or Cisco are commonly required.
More about Data Center Operations Engineer jobs

What cities are hiring for Data Center Operations Engineer jobs?

Cities with the most Data Center Operations Engineer job openings:

What states have the most Data Center Operations Engineer jobs?

States with the most job openings for Data Center Operations Engineer jobs include:

Infographic showing various Data Center Operations Engineer job openings in the United States as of August 2026, with employment types broken down into 84% Full Time, 13% Part Time, 1% Temporary, 1% Contract, and 1% Nights. Highlights an 94% Physical, 2% Hybrid, and 4% Remote job distribution, with an average salary of $147,461 per year, or $70.9 per hour.

Data Center Production Operations Engineer

Meta

Newark, CA • On-site

$111K/yr

Full-time

Re-posted 16 days ago


Key responsibilities

  • Manage and maintain large-scale server fleets, including hardware triage, failure analysis, and coordinating repair and replacement workflows

  • Monitor production systems health using observability tooling and telemetry data to proactively identify and resolve infrastructure anomalies

  • Develop and refine operational runbooks, escalation procedures, and incident response playbooks specific to data center server environments


Meta rating

7.8

Company rating: 7.8 out of 10

Based on 45 frontline employees who took The Breakroom Quiz

139th of 247 rated software companies


Job description

Meta is seeking a Data Center Production Operations Engineer to support the reliability, efficiency, and scalability of our global data center infrastructure. In this role, you will be responsible for the day-to-day operational health of server fleets and production systems that underpin Meta's family of apps and services. You will work at the intersection of hardware lifecycle management, systems reliability, and operational process improvement, ensuring that production environments meet the demands of billions of users worldwide.
Data Center Production Operations Engineer Responsibilities:
  • Manage and maintain large-scale server fleets across data center environments, including hardware triage, failure analysis, and coordinating repair and replacement workflows
  • Monitor production systems health using observability tooling and telemetry data to proactively identify and resolve infrastructure anomalies before they impact service availability
  • Develop and refine operational runbooks, escalation procedures, and incident response playbooks specific to data center server environments
  • Collaborate with hardware engineering, network operations, and capacity planning teams to support server deployment, decommissioning, and lifecycle transitions
  • Analyze failure trends and operational data to identify systemic issues in server hardware or firmware, and drive root cause analysis and corrective action
  • Contribute to automation initiatives that reduce manual toil in server provisioning, health checks, and fleet management workflows, including leveraging AI-integrated tooling
  • Partner with cross-functional teams to evaluate and implement process improvements that increase operational efficiency and reduce mean time to resolution for production incidents
  • Communicate infrastructure status, incident timelines, and risk assessments to engineering and operations stakeholders through clear written and verbal updates
  • Support capacity readiness activities by validating server acceptance criteria and coordinating with data center technicians during hardware bring-up and commissioning
  • Identify gaps in monitoring coverage or operational tooling and propose solutions that improve fleet visibility and production reliability
  • Participate in 24/7 on-call rotation
  • Ability to travel up to 15% of the time

Minimum Qualifications:
  • 6+ years of experience in data center operations, site operations, or production infrastructure engineering supporting large-scale server environments
  • 6+ years of experience with server hardware components including CPUs, memory, storage, and network interface cards, including hands-on troubleshooting and failure diagnosis
  • Experience using systems monitoring and observability platforms to track fleet health, identify anomalies, and drive incident resolution in production data center environments
  • Experience developing or improving operational processes, runbooks, or automation scripts to support server fleet management at scale
  • Experience collaborating with hardware engineering, network, and capacity teams to coordinate infrastructure deployments and lifecycle activities

Preferred Qualifications:
  • Experience contributing to post-incident reviews and translating findings into durable operational improvements that reduce recurrence across a server fleet
  • Experience with scripting languages such as Python or Bash to automate data center operations tasks including health checks, inventory management, or alerting workflows
  • Familiarity with server firmware management, BIOS configuration, and out-of-band management interfaces such as IPMI or Redfish in hyperscale data center environments
  • Background in capacity planning or hardware acceptance testing processes within a large-scale cloud or hyperscale data center organization

About Meta:
Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today—beyond the constraints of screens, the limits of distance, and even the rules of physics.
Meta is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment.
Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at accommodations-ext@meta.com.
$111,010/year to $158,995/year + bonus + equity + benefits
Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate, monthly rate, or annual salary only, and do not include bonus, equity or sales incentives, if applicable. In addition to base compensation, Meta offers benefits. Learn more about benefits at Meta.

What Meta employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom