Network Engineer, AI Infrastructure Repair Responsibilities: * Define and drive the long-term ... Experience developing and driving strategy for network fault management, repair automation, or ...
Network Engineer, AI Infrastructure Repair Responsibilities: * Define and drive the long-term ... Experience developing and driving strategy for network fault management, repair automation, or ...
We conduct research and development, manage national laboratories, design and manufacture products ... Conduct hardware bring-up, electrical testing, debugging, and fault isolation using laboratory ...
We conduct research and development, manage national laboratories, design and manufacture products ... Conduct hardware bring-up, electrical testing, debugging, and fault isolation using laboratory ...
We conduct research and development, manage national laboratories, design and manufacture products ... Conduct hardware bring-up, electrical testing, debugging, and fault isolation using laboratory ...
We conduct research and development, manage national laboratories, design and manufacture products ... Conduct hardware bring-up, electrical testing, debugging, and fault isolation using laboratory ...
Manage and monitor spare parts inventory, equipment fault indicators, quality standards, and ... Bachelor's degree in mechanical engineering, Electrical Engineering, Industrial Engineering, or a ...
Manage and monitor spare parts inventory, equipment fault indicators, quality standards, and ... Bachelor's degree in mechanical engineering, Electrical Engineering, Industrial Engineering, or a ...
... files (fault_script variants) & .fault files) o Write BUILD files for fault file creation o ... Nothing in this restricts management's right to assign or reassign duties and responsibilities to ...
... files (fault_script variants) & .fault files) o Write BUILD files for fault file creation o ... Nothing in this restricts management's right to assign or reassign duties and responsibilities to ...
Fleet Reliability Specialist (NJUS)
Columbus, OH · On-site
$53.50 - $71.25/hr
... fault data within the Fault Data Management (FDM) process. This role serves as the primary ... Associate's or Bachelor's degree in Aviation, Engineering, Data Analytics, or related field ...
Fleet Reliability Specialist (NJUS)
Columbus, OH · On-site
$53.50 - $71.25/hr
... fault data within the Fault Data Management (FDM) process. This role serves as the primary ... Associate's or Bachelor's degree in Aviation, Engineering, Data Analytics, or related field ...
We are currently hiring an Equipment Engineer for our new state-of-the-art factory in the Columbus ... Manage and monitor spare parts inventory, equipment fault indicators, quality standards, and ...
We are currently hiring an Equipment Engineer for our new state-of-the-art factory in the Columbus ... Manage and monitor spare parts inventory, equipment fault indicators, quality standards, and ...
Co-manage Fibre Channel, VxLAN, and NAS environments with a continued focus on: * Fault tolerance ... Serve as an engineering-level escalation resource for the Network Operations Center (NOC) and ...
Co-manage Fibre Channel, VxLAN, and NAS environments with a continued focus on: * Fault tolerance ... Serve as an engineering-level escalation resource for the Network Operations Center (NOC) and ...
Mainframe Site Reliability Engineering (SRE)
Columbus, OH · On-site
$55 - $73.25/hr
Lead incident management, root cause analysis (RCA), and post-incident reviews * Partner with ... Experience with resilience engineering, chaos testing, or fault injection concepts * Prior people ...
Quick apply
Mainframe Site Reliability Engineering (SRE)
Columbus, OH · On-site
$55 - $73.25/hr
Lead incident management, root cause analysis (RCA), and post-incident reviews * Partner with ... Experience with resilience engineering, chaos testing, or fault injection concepts * Prior people ...
Configuration of real-time controllers to perform custom logic such as automatic system fault ... Technical leading to assist Project Managers and Salesforce in planning applications for customers
Configuration of real-time controllers to perform custom logic such as automatic system fault ... Technical leading to assist Project Managers and Salesforce in planning applications for customers
The Senior Infrastructure Engineer applies their deep background designing, managing, and building ... Design of self-healing and fault-tolerant services * Familiarity developing with RESTful API ...
The Senior Infrastructure Engineer applies their deep background designing, managing, and building ... Design of self-healing and fault-tolerant services * Familiarity developing with RESTful API ...
The Senior Infrastructure Engineer applies their deep background designing, managing, and building ... Design of self-healing and fault-tolerant services * Familiarity developing with RESTful API ...
The Senior Infrastructure Engineer applies their deep background designing, managing, and building ... Design of self-healing and fault-tolerant services * Familiarity developing with RESTful API ...
Senior Electrical Engineer - Mission Critical
Columbus, OH · On-site
$103K - $135K/yr
Conducts fault studies and load flow analysis. Develop design criteria for protective device ... Provides electrical design direction and resource management on multi-discipline projects.
Senior Electrical Engineer - Mission Critical
Columbus, OH · On-site
$103K - $135K/yr
Conducts fault studies and load flow analysis. Develop design criteria for protective device ... Provides electrical design direction and resource management on multi-discipline projects.
Electrical Systems Engineer
Delaware, OH · On-site
Key activities include launching new products and services, managing the existing portfolio of ... and withstand fault performance of the STS and PDU units. * Prepare detailed technical ...
Electrical Systems Engineer
Delaware, OH · On-site
Key activities include launching new products and services, managing the existing portfolio of ... and withstand fault performance of the STS and PDU units. * Prepare detailed technical ...
Electrical Systems Engineer
Delaware, OH · On-site
Key activities include launching new products and services, managing the existing portfolio of ... and withstand fault performance of the STS and PDU units. * Prepare detailed technical ...
Electrical Systems Engineer
Delaware, OH · On-site
Key activities include launching new products and services, managing the existing portfolio of ... and withstand fault performance of the STS and PDU units. * Prepare detailed technical ...
Key activities include launching new products and services, managing the existing portfolio of ... and withstand fault performance of the STS and PDU units. * Prepare detailed technical ...
Key activities include launching new products and services, managing the existing portfolio of ... and withstand fault performance of the STS and PDU units. * Prepare detailed technical ...
Key activities include launching new products and services, managing the existing portfolio of ... and withstand fault performance of the STS and PDU units. * Prepare detailed technical ...
Key activities include launching new products and services, managing the existing portfolio of ... and withstand fault performance of the STS and PDU units. * Prepare detailed technical ...
Key activities include launching new products and services, managing the existing portfolio of ... performance, and withstand fault performanceof the STS and PDU units.
Key activities include launching new products and services, managing the existing portfolio of ... performance, and withstand fault performanceof the STS and PDU units.
Electrical Systems Engineer
Delaware, OH · On-site
Key activities include launching new products and services, managing the existing portfolio of ... and withstand fault performance of the STS and PDU units. * Prepare detailed technical ...
Electrical Systems Engineer
Delaware, OH · On-site
Key activities include launching new products and services, managing the existing portfolio of ... and withstand fault performance of the STS and PDU units. * Prepare detailed technical ...
Sr. Electrical Engineer - Power
$102K - $132K/yr
Define and manage interfaces between electrical, mechanical, thermal, and control subsystems ... Perform short-circuit, fault, and coordination analysis * Select and specify electrical components ...
Sr. Electrical Engineer - Power
$102K - $132K/yr
Define and manage interfaces between electrical, mechanical, thermal, and control subsystems ... Perform short-circuit, fault, and coordination analysis * Select and specify electrical components ...
Fault Management Engineer information
See salary details
$29.5K - $43.5K
6% of jobs
$43.5K - $57.5K
9% of jobs
$57.5K - $71.5K
7% of jobs
$75.4K is the 25th percentile. Wages below this are outliers.
$71.5K - $85.5K
10% of jobs
$85.5K - $99.5K
10% of jobs
The median wage is $108.8K / yr.
$99.5K - $113.5K
13% of jobs
$113.5K - $127.5K
13% of jobs
$138K is the 75th percentile. Wages above this are outliers.
$127.5K - $141.5K
11% of jobs
$141.5K - $155.5K
9% of jobs
$155.5K - $169.5K
7% of jobs
$169.5K - $183.5K
6% of jobs
$29.5K
$111.1K
$183.5K
How much do fault management engineer jobs pay per year?

Meta rating
7.8
Based on 45 frontline employees who took The Breakroom Quiz
136th of 246 rated software companies
Job description
Network Engineer, AI Infrastructure Repair Responsibilities:
- Define and drive the long-term strategy for AI network repair and remediation programs across large-scale data center environments supporting machine learning workloads
- Lead root cause analysis and resolution of complex network faults affecting high-performance AI training and inference fabrics, including RDMA, high-speed Ethernet, and optical interconnect layers
- Develop and champion novel approaches to network fault detection, automated remediation, and repair workflow optimization for AI cluster infrastructure
- Partner with hardware, software, and data center operations teams to align network repair programs with AI infrastructure deployment roadmaps and capacity plans
- Establish and refine operational frameworks, runbooks, and tooling for network repair at scale, reducing mean time to repair across AI fabric environments
- Identify systemic reliability risks in AI network infrastructure and drive cross-functional initiatives to address them before they impact production workloads
- Influence the design of next-generation AI network architectures by contributing repair and reliability insights to hardware and topology decisions
- Leverage AI-driven analytics and automation tools to redesign repair workflows, accelerating fault identification and resolution across distributed network environments
- Build and maintain strategic relationships with internal engineering, operations, and vendor partners to ensure repair programs scale with AI infrastructure growth
- Communicate program status, risk, and strategic recommendations to engineering leaders and cross-functional stakeholders through structured reporting and executive briefings
Minimum Qualifications:
- Experience influencing technical direction and organizational strategy through data-driven analysis, written proposals, and stakeholder alignment across engineering and operations teams
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
- Experience leading cross-functional programs that span network operations, hardware deployment, and infrastructure reliability at data center scale
- Experience developing and driving strategy for network fault management, repair automation, or remediation programs in production environments
- Experience designing, deploying, or operating high-speed network fabrics used in AI or machine learning infrastructure, including technologies such as RDMA over Converged Ethernet, InfiniBand, or high-density optical interconnects
- 12+ years of experience in network engineering, with a focus on large-scale data center or high-performance computing network environments
Preferred Qualifications:
- Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
- Experience with network telemetry platforms, observability tooling, or AI-assisted anomaly detection applied to large-scale fabric environments
- Experience building or scaling repair operations programs, including workforce planning, tooling development, and process standardization across multiple data center sites
- Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
- Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
- Track record of contributing to network hardware or topology design reviews, translating operational repair insights into upstream engineering improvements
- Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
- Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
- Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
- Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
- Familiarity with AI accelerator interconnect architectures and the network reliability requirements of distributed training workloads at hyperscale
About Meta:
Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today—beyond the constraints of screens, the limits of distance, and even the rules of physics.
Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at accommodations-ext@meta.com.
$193,000/year to $271,000/year + bonus + equity + benefits
Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate, monthly rate, or annual salary only, and do not include bonus, equity or sales incentives, if applicable. In addition to base compensation, Meta offers benefits. Learn more about benefits at Meta.
About Meta
Sourced by ZipRecruiter
Industry
Internet and it, media and telecom and software development
Company size
10,000+ Employees
Headquarters location
Menlo Park, CA, US