Meta is building the next generation of AI infrastructure to power large-scale machine learning workloads, and the reliability of that infrastructure depends on reliable, high-performance network ...
Meta is building the next generation of AI infrastructure to power large-scale machine learning workloads, and the reliability of that infrastructure depends on reliable, high-performance network ...
Machine Learning Engineer II
$94K - $128K/yr
Machine Learning II Engineer - Incydr Product Development Mimecast is at the forefront of the ... Show capability in developing and maintaining infrastructure-as-code to automate and streamline the ...
Machine Learning Engineer II
$94K - $128K/yr
Machine Learning II Engineer - Incydr Product Development Mimecast is at the forefront of the ... Show capability in developing and maintaining infrastructure-as-code to automate and streamline the ...
Machine Learning Engineer II
Columbus, OH · On-site
$94K - $128K/yr
Machine Learning II Engineer - Incydr Product Development Mimecast is at the forefront of the ... Show capability in developing and maintaining infrastructure-as-code to automate and streamline the ...
Machine Learning Engineer II
Columbus, OH · On-site
$94K - $128K/yr
Machine Learning II Engineer - Incydr Product Development Mimecast is at the forefront of the ... Show capability in developing and maintaining infrastructure-as-code to automate and streamline the ...
Machine Learning Engineer
Columbus, OH · On-site
... and building the infrastructure that makes world-scale RL training possible. This is a high ... deep learning proficiency (PyTorch preferred; familiar with training loops, optimizers, mixed ...
Machine Learning Engineer
Columbus, OH · On-site
... and building the infrastructure that makes world-scale RL training possible. This is a high ... deep learning proficiency (PyTorch preferred; familiar with training loops, optimizers, mixed ...
Senior Machine Learning Engineer
Columbus, OH · On-site +1
$100K - $138K/yr
We are seeking a Senior Machine Learning Engineer to lead the development of a neural welding ... Partner with controls, welding, robotics, world-model, data, and ML infrastructure engineers.
Senior Machine Learning Engineer
Columbus, OH · On-site +1
$100K - $138K/yr
We are seeking a Senior Machine Learning Engineer to lead the development of a neural welding ... Partner with controls, welding, robotics, world-model, data, and ML infrastructure engineers.
Sr AI Machine Learning Engineer
Columbus, OH · Hybrid
$117K - $175K/yr
The Hartford is seeking Senior AI Machine Learning Engineer to build Machine Learning Operations ... Experience with IAC (Infrastructure as Code) including Cloud Formation, Terraform, or equivalents
Sr AI Machine Learning Engineer
Columbus, OH · Hybrid
$117K - $175K/yr
The Hartford is seeking Senior AI Machine Learning Engineer to build Machine Learning Operations ... Experience with IAC (Infrastructure as Code) including Cloud Formation, Terraform, or equivalents
We are seeking a Machine Learning Engineer to join us as a founding member. You will be among the ... Stand up ML infrastructure - training pipelines, experiment tracking, data versioning, reproducible ...
We are seeking a Machine Learning Engineer to join us as a founding member. You will be among the ... Stand up ML infrastructure - training pipelines, experiment tracking, data versioning, reproducible ...
We are seeking a Machine Learning Engineer to join us as a founding member. You will be among the ... Stand up ML infrastructure - training pipelines, experiment tracking, data versioning, reproducible ...
We are seeking a Machine Learning Engineer to join us as a founding member. You will be among the ... Stand up ML infrastructure - training pipelines, experiment tracking, data versioning, reproducible ...
Meta is building the next generation of AI infrastructure to power large-scale machine learning workloads, and the reliability of that infrastructure de.
Meta is building the next generation of AI infrastructure to power large-scale machine learning workloads, and the reliability of that infrastructure de.
Work you'll do As a Cloud Security Senior Manager, Azure Infrastructure & AI on the Enterprise ... Building and automating secure cloud and machine learning operations (MLOps) pipelines using tools ...
Work you'll do As a Cloud Security Senior Manager, Azure Infrastructure & AI on the Enterprise ... Building and automating secure cloud and machine learning operations (MLOps) pipelines using tools ...
Infrastructure Engineer Staff - Netcool (NOI) SME
Gahanna, OH · On-site
$101K - $132K/yr
Machine learning-enabled monitoring Preferred Skills * Utility industry monitoring experience * SOX ... Infrastructure monitoring design * Enterprise observability practices * Network monitoring and ...
Infrastructure Engineer Staff - Netcool (NOI) SME
Gahanna, OH · On-site
$101K - $132K/yr
Machine learning-enabled monitoring Preferred Skills * Utility industry monitoring experience * SOX ... Infrastructure monitoring design * Enterprise observability practices * Network monitoring and ...
Infrastructure Engineer Staff - Netcool (NOI) SME
$101K - $132K/yr
Machine learning-enabled monitoring Preferred Skills * Utility industry monitoring experience * SOX ... Infrastructure monitoring design * Enterprise observability practices * Network monitoring and ...
Infrastructure Engineer Staff - Netcool (NOI) SME
$101K - $132K/yr
Machine learning-enabled monitoring Preferred Skills * Utility industry monitoring experience * SOX ... Infrastructure monitoring design * Enterprise observability practices * Network monitoring and ...
Senior Software Engineer, Machine Learning
Columbus, OH · On-site
$118K - $156K/yr
You'll work closely with our research teams and our data platform team, leveraging their data and ML infrastructure expertise, while focusing on creating a robust, scalable, and developer-friendly ...
Senior Software Engineer, Machine Learning
Columbus, OH · On-site
$118K - $156K/yr
You'll work closely with our research teams and our data platform team, leveraging their data and ML infrastructure expertise, while focusing on creating a robust, scalable, and developer-friendly ...
Senior Software Engineer, Machine Learning
$118K - $156K/yr
You'll work closely with our research teams and our data platform team, leveraging their data and ML infrastructure expertise, while focusing on creating a robust, scalable, and developer-friendly ...
Senior Software Engineer, Machine Learning
$118K - $156K/yr
You'll work closely with our research teams and our data platform team, leveraging their data and ML infrastructure expertise, while focusing on creating a robust, scalable, and developer-friendly ...
This role demands deep hands-on expertise with Databricks Machine Learning, MLflow, and a broad ... Implement Infrastructure as Code using Terraform to provision and manage Databricks and AWS ...
This role demands deep hands-on expertise with Databricks Machine Learning, MLflow, and a broad ... Implement Infrastructure as Code using Terraform to provision and manage Databricks and AWS ...
ML Engineer
Columbus, OH · On-site
Build production-ready machine learning systems and infrastructure. * Productionize machine learning models developed by Data Science teams. * Design and deploy Large Language Model (LLM ...
ML Engineer
Columbus, OH · On-site
Build production-ready machine learning systems and infrastructure. * Productionize machine learning models developed by Data Science teams. * Design and deploy Large Language Model (LLM ...
Applied AI ML-Vice President
Columbus, OH · On-site
... in machine learning frameworks, ML Ops tools and practices. • Strong proficiency in engineering programming languages (e.g., Python, Java) and infrastructure as code (e.g., Terraform ...
Applied AI ML-Vice President
Columbus, OH · On-site
... in machine learning frameworks, ML Ops tools and practices. • Strong proficiency in engineering programming languages (e.g., Python, Java) and infrastructure as code (e.g., Terraform ...
... data infrastructure and supports the senior leadership with insights, management reports, and ... Conducts research on cutting-edge techniques and tools in machine learning/deep learning/artificial ...
... data infrastructure and supports the senior leadership with insights, management reports, and ... Conducts research on cutting-edge techniques and tools in machine learning/deep learning/artificial ...
Senior Data Scientist - AI
Columbus, OH · On-site +1
... data infrastructure and supports the senior leadership with insights, management reports, and ... Conducts research on cutting-edge techniques and tools in machine learning/deep learning/artificial ...
Senior Data Scientist - AI
Columbus, OH · On-site +1
... data infrastructure and supports the senior leadership with insights, management reports, and ... Conducts research on cutting-edge techniques and tools in machine learning/deep learning/artificial ...
... data infrastructure and supports the senior leadership with insights, management reports, and ... Conducts research on cutting-edge techniques and tools in machine learning/deep learning/artificial ...
... data infrastructure and supports the senior leadership with insights, management reports, and ... Conducts research on cutting-edge techniques and tools in machine learning/deep learning/artificial ...
Machine Learning Infrastructure information
See salary details
$15.87 - $19.23
6% of jobs
$21.59 is the 25th percentile. Wages below this are outliers.
$19.23 - $22.60
27% of jobs
The median wage is $24.67 / hr.
$22.60 - $25.96
28% of jobs
$25.96 - $29.33
14% of jobs
$29.45 is the 75th percentile. Wages above this are outliers.
$29.33 - $32.69
15% of jobs
$32.69 - $36.06
6% of jobs
$36.06 - $39.42
2% of jobs
$39.42 - $42.79
1% of jobs
$42.79 - $46.15
1% of jobs
$46.15 - $49.52
0% of jobs
$49.52 - $52.88
0% of jobs
$15
$28
$52
How much do machine learning infrastructure jobs pay per hour?

Meta rating
7.8
Based on 45 frontline employees who took The Breakroom Quiz
136th of 246 rated software companies
Job description
Network Engineer, AI Infrastructure Repair Responsibilities:
- Define and drive the long-term strategy for AI network repair and remediation programs across large-scale data center environments supporting machine learning workloads
- Lead root cause analysis and resolution of complex network faults affecting high-performance AI training and inference fabrics, including RDMA, high-speed Ethernet, and optical interconnect layers
- Develop and champion novel approaches to network fault detection, automated remediation, and repair workflow optimization for AI cluster infrastructure
- Partner with hardware, software, and data center operations teams to align network repair programs with AI infrastructure deployment roadmaps and capacity plans
- Establish and refine operational frameworks, runbooks, and tooling for network repair at scale, reducing mean time to repair across AI fabric environments
- Identify systemic reliability risks in AI network infrastructure and drive cross-functional initiatives to address them before they impact production workloads
- Influence the design of next-generation AI network architectures by contributing repair and reliability insights to hardware and topology decisions
- Leverage AI-driven analytics and automation tools to redesign repair workflows, accelerating fault identification and resolution across distributed network environments
- Build and maintain strategic relationships with internal engineering, operations, and vendor partners to ensure repair programs scale with AI infrastructure growth
- Communicate program status, risk, and strategic recommendations to engineering leaders and cross-functional stakeholders through structured reporting and executive briefings
Minimum Qualifications:
- Experience influencing technical direction and organizational strategy through data-driven analysis, written proposals, and stakeholder alignment across engineering and operations teams
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
- Experience leading cross-functional programs that span network operations, hardware deployment, and infrastructure reliability at data center scale
- Experience developing and driving strategy for network fault management, repair automation, or remediation programs in production environments
- Experience designing, deploying, or operating high-speed network fabrics used in AI or machine learning infrastructure, including technologies such as RDMA over Converged Ethernet, InfiniBand, or high-density optical interconnects
- 12+ years of experience in network engineering, with a focus on large-scale data center or high-performance computing network environments
Preferred Qualifications:
- Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
- Experience with network telemetry platforms, observability tooling, or AI-assisted anomaly detection applied to large-scale fabric environments
- Experience building or scaling repair operations programs, including workforce planning, tooling development, and process standardization across multiple data center sites
- Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
- Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
- Track record of contributing to network hardware or topology design reviews, translating operational repair insights into upstream engineering improvements
- Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
- Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
- Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
- Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
- Familiarity with AI accelerator interconnect architectures and the network reliability requirements of distributed training workloads at hyperscale
About Meta:
Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today—beyond the constraints of screens, the limits of distance, and even the rules of physics.
Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at accommodations-ext@meta.com.
$193,000/year to $271,000/year + bonus + equity + benefits
Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate, monthly rate, or annual salary only, and do not include bonus, equity or sales incentives, if applicable. In addition to base compensation, Meta offers benefits. Learn more about benefits at Meta.
About Meta
Sourced by ZipRecruiter
Industry
Internet and it, media and telecom and software development
Company size
10,000+ Employees
Headquarters location
Menlo Park, CA, US