1

Datacenter Operations Jobs (NOW HIRING)

Manages Campus operations of the 24x7 Datacenter environment. Identifies and implements critical metrics necessary to manage and monitor the performance of overall operations including, but not ...

The Datacenter Operations team is dedicated to ensuring the seamless operation and management of the data center facilities. This team operates in multiple locations, including Phoenix and Reno, and ...

Experience with Hands and feed support for Data Center operations is a plus. * Candidates must be available to work onsite on a Rotating / Flexible shift timings at respective Datacenter to provide.

DC Devops Engineer

Dallas, TX · On-site

$52.25 - $71.50/hr

THE ROLE As part of a Data Center DevOps team, you will work in a local datacenter interfacing directly with the datacenter operations team to assist in providing post hardware repair support to ...

Datacenter Technician Location: Colorado City, TX Company Overview: Our client is a leading ... operation of our mission-critical computing resources. Key Responsibilities: - Perform regular ...

Datacenter Technician

Barker, NY · On-site

$25 - $32/hr

Datacenter Technician Location: Barker, NY Company Overview: Our client is a leading technology ... operation of our mission-critical computing resources. Key Responsibilities: - Perform regular ...

DC Operations - NYC

Manhattan, NY · On-site

$150K - $225K/mo

The Role We are seeking an experienced Datacentre Operations Engineer to own the physical layer of our client's Americas datacentre estate. You will lead colocation deployments, manage third-party ...

next page

Showing results 1-20

Datacenter Operations information

See salary details

$52K

$128.5K

$200K

How much do datacenter operations jobs pay per year?

As of Jul 20, 2026, the average yearly pay for datacenter operations in the United States is $128,526.00, according to ZipRecruiter salary data. Most workers in this role earn between $94,000.00 and $163,500.00 per year, depending on experience, location, and employer.

What is the difference between Datacenter Operations vs Data Center Technician?

AspectDatacenter OperationsData Center Technician
CertificationsCompTIA Server+, Cisco CCNA, Data Center certificationsCompTIA Server+, Cisco CCNA, Data Center certifications
Work EnvironmentData centers, server rooms, 24/7 operationsData centers, server rooms, hardware troubleshooting
Job FocusMonitoring, managing infrastructure, ensuring uptimeInstalling, maintaining, troubleshooting hardware
Employer & Industry UsageData center providers, IT departmentsData center providers, IT support teams

Both roles operate within data centers and require similar certifications. However, Datacenter Operations focuses on managing and monitoring infrastructure to ensure continuous uptime, while Data Center Technicians primarily handle hardware installation and troubleshooting. Understanding these differences helps in choosing the right career path or job search focus.

What are Datacenter Operations?

Datacenter Operations refer to the processes, activities, and staff responsible for managing and maintaining the technical infrastructure within a datacenter. This includes overseeing servers, storage, networking hardware, power, cooling systems, and security to ensure continuous uptime and optimal performance. Datacenter operations teams are also tasked with monitoring systems, troubleshooting issues, performing regular maintenance, and implementing disaster recovery plans. Their work is crucial to ensure that organizational data and IT services are available, secure, and reliable.

What are some common challenges faced in a Datacenter Operations role, and how can they be effectively managed?

Professionals in Datacenter Operations often encounter challenges such as maintaining high availability, managing physical and cybersecurity risks, and responding to hardware failures or outages quickly. Effective management involves following strict protocols, proactive monitoring, and collaborating closely with IT, network, and facilities teams to ensure seamless operations. Additionally, staying up-to-date with evolving technologies and adopting automation tools can help address these challenges and improve overall efficiency.

What are the key skills and qualifications needed to thrive in Datacenter Operations, and why are they important?

To thrive in Datacenter Operations, you need a solid understanding of IT infrastructure, hardware troubleshooting, and network management, typically supported by relevant certifications such as CompTIA Server+, Network+, or Cisco CCNA. Familiarity with data center management tools, monitoring systems, and ticketing platforms is commonly required. Strong problem-solving skills, attention to detail, and effective communication are essential soft skills for this role. These abilities ensure system reliability, efficient issue resolution, and seamless collaboration in maintaining critical business operations.
More about Datacenter Operations jobs
What cities are hiring for Datacenter Operations jobs? Cities with the most Datacenter Operations job openings:
What are the most commonly searched types of Datacenter Operations jobs? The most popular types of Datacenter Operations jobs are:
What states have the most Datacenter Operations jobs? States with the most job openings for Datacenter Operations jobs include:

Datacenter Operations Manager

WeEngage Group | B Corp™

Santa Rosa, CA • On-site

$182K/yr

Other

Posted 4 days ago


Job description

Data Center Site Lead – AI Infrastructure (Datacenter Operations Manager)


Location: Santa Clara, California


Working model: Full-time, on-site


Employment type: Permanent


About the opportunity

We are supporting a fast-growing AI infrastructure company that designs, deploys, and operates large-scale GPU compute environments.

The company is expanding its data center operations in Santa Clara and is looking for a hands-on Data Center Site Lead to take ownership of the site’s day-to-day operation, technical reliability, and future growth.

This is not a purely managerial position. You will be expected to understand the site in detail, including its power, cooling, networking, server infrastructure, dependencies, capacity constraints, and operational risks. You will act as the senior technical presence on-site, lead other technicians and vendors, and take ownership when incidents or equipment failures occur.

The environment supports demanding AI and high-performance computing workloads where uptime, response speed, and disciplined execution are critical.


Key responsibilities

  • Take day-to-day operational ownership of the Santa Clara data center site.
  • Act as the senior technical lead for on-site technicians, contractors, vendors, and remote-hands teams.
  • Install, configure, troubleshoot, and maintain GPU servers, storage systems, networking equipment, cabling, and supporting infrastructure.
  • Monitor site conditions, including power, cooling, temperature, humidity, capacity, alarms, and equipment health.
  • Ensure the availability and reliability of the site within a 24/7 operational environment.
  • Lead the response to hardware failures, environmental alarms, connectivity issues, and other critical incidents.
  • Own incident reporting from initial detection through root-cause analysis, corrective action, and final closure.
  • Track equipment downtime, identify recurring failure patterns, and introduce preventive measures.
  • Coordinate escalations with hardware manufacturers, colocation providers, network teams, and other technical partners.
  • Plan and schedule server installations, rack deployments, maintenance activities, upgrades, and hardware refreshes.
  • Support data hall expansions, cluster deployments, migrations, and new capacity coming online.
  • Maintain accurate records covering site assets, installations, incidents, maintenance work, capacity, and operational risks.
  • Ensure all work follows the company’s safety, security, access-control, and change-management procedures.
  • Help develop site operating procedures, escalation processes, maintenance schedules, and reliability standards.
  • Mentor junior technicians and support the recruitment and development of the on-site team as the facility grows.
  • Provide regular updates to leadership on uptime, incidents, staffing, capacity, operational risks, and planned work.
  • Participate in an on-call rotation and provide escalation support for critical incidents outside normal working hours.


What we are looking for

  • At least five years of experience within data center operations, critical infrastructure, cloud infrastructure, or another mission-critical technical environment.
  • Strong hands-on experience installing and supporting enterprise servers, storage, networking hardware, and structured cabling.
  • Previous experience acting as a site lead, senior technician, shift lead, operations manager, or technical escalation point.
  • Practical understanding of data center power, cooling, networking, rack layouts, environmental monitoring, and common infrastructure failure modes.
  • Experience managing incidents in a structured manner, including escalation, root-cause analysis, documentation, and preventive action.
  • Ability to work independently and make sound operational decisions without requiring constant supervision.
  • Experience coordinating technicians, contractors, vendors, and remote engineering teams.
  • Strong planning and organisational skills, particularly around installations, maintenance windows, upgrades, and capacity expansion.
  • Clear written and verbal communication skills.
  • Willingness to work on-site full-time in Santa Clara and participate in on-call coverage.


Particularly relevant experience

  • Supporting enterprise GPU platforms using NVIDIA Ampere, Hopper, Blackwell, GB200, GB300, or similar systems.
  • Operating high-density AI, HPC, hyperscale, or cloud infrastructure.
  • Direct liquid cooling, coolant distribution units, liquid loops, or other advanced cooling technologies.
  • Large GPU cluster deployments, server bring-up, burn-in, firmware updates, and hardware validation.
  • Data center build-outs, new data hall openings, migrations, expansions, or infrastructure refresh programmes.
  • DCIM, CMMS, monitoring, alerting, ticketing, and maintenance-management platforms.
  • Managing 24/7 shift coverage or supporting teams operating across multiple shifts.
  • Working within environments governed by strict SLAs, security controls, and safety procedures.
  • Data center qualifications such as CDCP, CDCS, or an equivalent certification.