1

Lte System Performance Engineer Jobs (NOW HIRING)

As an AI/HPC System Performance Engineer on the Network Infrastructure Engineering team, you will drive end-to-end performance characterization, bottleneck analysis, and optimization of large-scale ...

Alpharetta GA - Onsite Perform load tests to validate system performance and stability. Analyze tests results and work with Developers and Engineers to perform bug fixes. Provide technical assistance ...

This role supports propulsion system performance analysis and validation by collecting, formatting, analyzing, and reporting test performance data across development, acceptance, and qualification ...

The main responsibilities include writing and executing automation drivers for load tests, performing load tests to validate system performance, and collaborating with developers to resolve ...

Showing results 21-40

Lte System Performance Engineer information

See salary details

$11

$60

$98

How much do lte system performance engineer jobs pay per hour?

As of Sep 14, 2026, the average hourly pay for lte system performance engineer in the United States is $60.11, according to ZipRecruiter salary data. Most workers in this role earn between $49.28 and $68.03 per hour, depending on experience, location, and employer.

What are popular job titles related to Lte System Performance Engineer jobs?

For Lte System Performance Engineer jobs, the most frequently searched job titles are:

Infographic showing various Lte System Performance Engineer job openings in the United States as of August 2026, with employment types broken down into 1% As Needed, 88% Full Time, 9% Part Time, and 2% Contract. Highlights an 85% Physical, 2% Hybrid, and 13% Remote job distribution, with an average salary of $125,019 per year, or $60.1 per hour.

AI/HPC System Performance Engineer

Menlo Park, CA • On-site

Meta
Internet and IT • 10K+ employees

$154K - $217K/yr

Full-time

Re-posted yesterday


Key responsibilities

  • Profile and benchmark AI training and inference workloads across large-scale HPC clusters to identify network, compute, and memory bottlenecks

  • Develop and maintain performance analysis frameworks and dashboards to track system-level metrics including GPU utilization, network bandwidth, latency, and collective communication efficiency

  • Investigate and resolve performance regressions in distributed AI training environments, including issues related to RDMA fabrics, collective communication libraries, and job scheduling


Meta rating

7.8

Company rating: 7.8 out of 10

Based on 45 frontline employees who took The Breakroom Quiz


Job description

Meta is building large-scale AI and high-performance computing infrastructure to power next-generation AI research and products. As an AI/HPC System Performance Engineer on the Network Infrastructure Engineering team, you will drive end-to-end performance characterization, bottleneck analysis, and optimization of large-scale AI training and inference clusters. In this role, you will work at the intersection of network fabric design, distributed computing, and AI workload behavior to ensure Meta's HPC systems deliver maximum throughput and efficiency for frontier model development.
Responsibilities
Profile and benchmark AI training and inference workloads across large-scale HPC clusters to identify network, compute, and memory bottlenecks
• Develop and maintain performance analysis frameworks and dashboards to track system-level metrics including GPU utilization, network bandwidth, latency, and collective communication efficiency
• Investigate and resolve performance regressions in distributed AI training environments, including issues related to RDMA fabrics, collective communication libraries, and job scheduling
• Collaborate with network infrastructure, hardware, and AI research teams to define performance requirements and validate new HPC cluster configurations
• Design and execute capacity and scalability experiments to inform network topology decisions for AI supercomputing infrastructure
• Build tooling and automation to continuously monitor HPC system health, detect anomalies, and reduce mean time to mitigation during performance incidents
• Establish service level objectives for AI cluster network performance and drive cross-functional alignment on reliability and efficiency targets
• Lead technical design reviews for network and system architecture changes affecting AI workload performance, communicating trade-offs clearly to engineering and product stakeholders
• Mentor other engineers on HPC performance methodologies, debugging techniques, and instrumentation best practices
• Leverage AI-assisted workflows to accelerate root cause analysis, automate routine performance reporting, and expand coverage across the HPC stack
Minimum Qualifications
• 5+ years of coding experience in C, C++, Python, or similar programming languages. Flexible to learn new programming languages
• Experience profiling and optimizing distributed AI or HPC workloads, including familiarity with GPU interconnects, RDMA networking, and collective communication frameworks such as NCCL or MPI
• Experience debugging complex, non-reproducible performance issues across multi-layer systems including network fabric, operating system, and application layers
• Experience designing and implementing performance monitoring systems, including instrumentation, telemetry pipelines, and alerting for large-scale infrastructure
• Experience driving cross-functional technical projects from requirements definition through production deployment, including communicating performance findings and trade-offs to diverse stakeholders
• Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
• 6+ years of experience in system performance engineering, network infrastructure engineering, or a related field within large-scale distributed computing or HPC environments
Preferred Qualifications
• Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
• Understanding of RDMA congestion control mechanisms on IB and RoCE Networks
• Understanding of AI training workloads and demands they exert on networks
• Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
• Understanding of the latest artificial intelligence (AI) technologies
• Experience in developing systems software in languages like C++
• Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
• Experience with machine learning frameworks such as PyTorch and TensorFlow
About Meta
Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today-beyond the constraints of screens, the limits of distance, and even the rules of physics.
Equal Employment Opportunity
Meta is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or other applicable legally protected characteristics. You may view our Equal Employment Opportunity notice here.

What Meta employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom