1

Infiniband Jobs in Tennessee (NOW HIRING)

Experience with InfiniBand networks and diagnostics. * Extensive experience with High Performance Parallel File Systems (Lustre, WEKA, GPFS, etc). * Experience with performance and diagnostic tools ...

RDMA, DPUs, Infiniband, many-core CPUs * Experience with declarative CI/CD tools such as ArgoCD * Experience with workflow engines such as Apache Airflow or Argo Workflows * Experience with ...

Partner with Observability and Telemetry teams to improve visibility into GPU health, network fabric performance, RDMA, RoCE, InfiniBand, and overall service health. * Collaborate with internal ...

Experience with Infiniband networks and diagnostics. * Extensive experience with High Performance Parallel File Systems (Lustre, WEKA, GPFS, etc). * Experience with performance and diagnostic tools ...

Senior Network Operations Engineer

Nashville, TN · On-site

$100K - $137K/yr

Juniper, Cisco, Arista, InfiniBand, NVIDIA, firewalls, routers, switches, circuit management, and optical/network transport services. * Strong analytical skills, including the ability to gather ...

Senior Network Operations Engineer

Nashville, TN · On-site

$100K - $137K/yr

Juniper, Cisco, Arista, InfiniBand, NVIDIA, firewalls, routers, switches, circuit management, and optical/network transport services. * Strong analytical skills, including the ability to gather ...

next page

Showing results 1-20

Infiniband information

What is InfiniBand?

Infiniband is a high-speed, low-latency networking technology commonly used in data centers and high-performance computing environments. It is designed to connect servers, storage systems, and network devices, providing much faster data transfer rates than traditional Ethernet. Infiniband supports scalable bandwidth and efficient communication, which makes it ideal for applications requiring rapid data movement, such as scientific simulations and large-scale database transactions. Its architecture also supports remote direct memory access (RDMA), which further reduces latency and CPU overhead.

What are the typical responsibilities of an InfiniBand network engineer in a data center environment?

InfiniBand network engineers are primarily responsible for designing, deploying, and maintaining high-performance InfiniBand fabrics that connect servers and storage systems in data centers, especially in HPC (High-Performance Computing) environments. Their daily tasks include monitoring network performance, troubleshooting connectivity or latency issues, and performing firmware and driver updates on InfiniBand switches and host adapters. They also collaborate closely with system administrators and application teams to optimize throughput and ensure reliable, low-latency communication. Additionally, InfiniBand engineers often participate in capacity planning and help scale the network infrastructure to meet growing computational demands.

What are the key skills and qualifications needed to thrive as an InfiniBand network engineer, and why are they important?

To thrive as an InfiniBand Network Engineer, you need a strong background in computer networking, Linux system administration, and high-performance computing (HPC) environments, often supported by a degree in computer science or related field. Familiarity with InfiniBand architecture, experience with tools like OpenFabrics Enterprise Distribution (OFED), and certifications such as CompTIA Network+ are valuable. Strong problem-solving skills, attention to detail, and effective communication are crucial soft skills for this role. These abilities are essential for ensuring efficient, reliable InfiniBand network performance in complex HPC or data center environments.

What is the difference between Infiniband vs Ethernet Network Engineer?

AspectInfinibandEthernet Network Engineer
Required CredentialsNetworking certifications, Cisco, Cisco CCNA, CCNPNetworking certifications, Cisco, CCNA, CCNP
Work EnvironmentData centers, high-performance computing environmentsCorporate networks, data centers, enterprise environments
Industry UsageHigh-performance computing, research institutionsBusiness, telecommunications, enterprise IT
Common Search/ComparisonYesYes

Infiniband and Ethernet Network Engineers both work with network infrastructure, but Infiniband specializes in high-speed, low-latency connections used in data centers and HPC environments. Ethernet Network Engineers focus on standard Ethernet networks used across various industries. While their certifications and skills overlap, their work environments and applications differ significantly.

What are popular job titles related to Infiniband jobs in Tennessee?

For Infiniband jobs in Tennessee, the most frequently searched job titles are:

Infographic showing various Infiniband job openings in Tennessee as of August 2026, with employment types broken down into 96% Full Time, and 4% Contract. Highlights an 79% Physical, 4% Hybrid, and 17% Remote job distribution.

Senior Network Operations Engineer - Nashville TN

Nashville, TN • On-site

$100K - $137K/yr

Other

Medical, Dental, Vision, Life, Retirement, PTO

Posted 10 days ago


Job description

Senior Network Operations Engineer - Nashville TN Job Description

In this role, you will monitor and troubleshoot network events, collect and analyze technical data, triage and mitigate incidents, coordinate escalations, and help drive continuous operational improvement. You will work alongside experienced engineers, partner teams, and vendors to maintain and optimize the infrastructure that supports Oracle customers and services worldwide.

Our mission is to keep OCI’s global network highly available, performant, and resilient—while delivering exceptional service to customers and dependable operational support to our engineering and technical teams.

For GNOC engineers, that mission translates into a broad, fast-moving role with real operational impact. You will help centrally manage OCI’s network infrastructure, respond to and resolve complex events, and develop automated solutions that reduce recurring operational work and improve reliability at scale.

Highly preferred work location is Nashville, TN.

Responsibilities

NOC Operations

Use established procedures and operational tooling to plan, implement, and safely complete network changes.

Mentor, onboard, and train junior network engineers.

Participate in operational rotations and provide break/fix and incident-response support.

Identify and triage actionable incidents through monitoring systems; analyze and mitigate network events; conduct or support root‑cause analysis (RCA); and coordinate follow‑up actions with internal support teams and vendors.

Provide on‑call support as required, exercising sound independent judgment in a varied and complex operational environment.

Participate in major incident calls and use technical and analytical skills to resolve network issues affecting Oracle customers and services.

Manage fault detection, response, and escalation for OCI systems and networks, collaborating with third‑party suppliers through resolution.

Leadership

Collaborate with GNOC Shift Leads and management to ensure the efficient and timely completion of daily GNOC responsibilities.

Lead, contribute to, and participate in the identification, development, and evaluation of projects and tools that improve GNOC effectiveness.

Drive runbook audits and updates to maintain compliance and align operational processes with partner service teams.

Conduct interviews and participate in hiring junior‑level engineers.

Lead and/or represent the GNOC in vendor meetings, service reviews, and governance boards.

Automation and Scripting

Collaborate with network automation teams to integrate and improve operational support tooling.

Develop scripts and automation to reduce manual effort and improve the reliability of routine operational tasks.

Preferred experience with Python, Puppet, SQL, Ansible, network automation, and databases.

Project Delivery

Lead technical initiatives, including the development and improvement of runbooks, methods of procedure (MOPs), operational processes, and team onboarding materials.

Support the implementation of short‑, medium‑, and long‑term plans to achieve project objectives.

Regularly engage senior management and network leadership to ensure team priorities and project objectives are met.

TECHNICAL QUALIFICATIONS

Networking

Strong knowledge of networking protocols and technologies, including BGP, OSPF, IS‑IS, TCP/IP, IPv4/IPv6, DNS, DHCP, MPLS, VPNs, and TLS.

Broad hands‑on experience with at least three of the following: Juniper, Cisco, Arista, InfiniBand, firewalls, routers, switches, circuit management, and optical/network transport services.

Strong analytical skills, including the ability to gather, correlate, and interpret data from multiple sources.

Ability to diagnose, prioritize, resolve, or appropriately elevate network alerts and faults.

Experience in a large ISP, cloud provider, or similarly complex enterprise network environment.

Exposure to commodity Ethernet hardware and networking ASICs, including Broadcom and NVIDIA/Mellanox.

Cisco, Arista and Juniper certifications are desirable.

GPU, RDMA, and HPC

Experience supporting GPU and RDMA network environments is highly desirable.

Experience supporting high‑performance computing (HPC) environments is highly desirable.

Experience with InfiniBand and NVIDIA networking technologies, including Spectrum, is highly desirable.

Network Design and Lifecycle Management

Participate in network lifecycle management, including network build, refresh, and upgrade projects.

Participate in network solution design and design‑review activities.

SOFT SKILLS AND OTHER DESIRED EXPERIENCE

Self‑motivated, proactive, and able to work independently.

Bachelor’s degree preferred, with 3–5 years of relevant network operations or engineering experience.

Strong organizational, time‑management, verbal, and written communication skills.

Comfortable managing a broad range of priorities in a fast‑paced operational environment.

Experience with incident‑response plans, processes, and strategies.

Experience supporting large‑scale enterprise infrastructure and cloud computing environments in a 24/7 network operations setting, including willingness to work rotational shifts.

Key Responsibilities
  • Takes proactive steps to design and architect infrastructure and/or service according to terms for reliability and functionality.
  • Forecasts demands for infrastructure and responds to capacity needs, ensuring systems have sufficient resources to handle current and future workloads.
  • Collaborates with the software development team to develop infrastructures and features that are reliable and scalable according to deployment requirements.
  • Independently identifies opportunities for and drives prototyping (e.g., testing new applications or infrastructures, assisting in onboarding).
  • Performs data collection, triage, technical analysis, and redirection to maintain and optimize operations and infrastructure reliability.
  • Independently monitors services, maintains up‑to‑date knowledge of their performance, and documents their condition.
  • Leverages comprehensive knowledge to perform incident response, root cause analyses, and/or maintenance on assigned services (e.g., software installs, version upgrades, security updates, backup and recovery).
  • Provides health and performance reporting and takes appropriate actions based on trends in data.
  • May independently perform provisioning to support infrastructure, applications, and services.
  • May perform standard and non‑standard decommissioning (e.g., shutting down servers, removing data from databases) to remove objects that are no longer needed.
  • Identifies opportunities for automation and assesses potential benefits.
  • Develops automation tools or scripts to provide solutions, gather metrics, monitor, analyze, mitigate, or remediate issues/defects within infrastructures.
  • Independently conducts testing to ensure automation performs the task correctly and produces expected results.
  • Communicates the scale, capacity, security, performance attributes, and requirements of services and technology within and sometimes beyond immediate team.
  • Identifies and explains the potential impact of infrastructure, feature, and tool changes, considering their impact on team operations.
  • Provides operational support for technology, escalating incidents and other standard and non‑standard issues arising within Oracle services.
  • Participates in on‑call shifts to address issues.
  • Resolves technical issues spanning various services, investigating and debugging products in order to reach SLOs (service level objectives).
  • Documents incidents and performs root cause analyses according to standard reporting methods.
  • Independently performs post‑mortem procedures to prevent incident reoccurrence.
  • Experiments with new tools and technologies to assess their potential impact on and improve infrastructure performance and reliability, ensuring adherence to security standards.
  • Independently identifies and executes improvements for performance bottlenecks and deployments to ensure efficient resource usage, speed, and scalability.
  • Develops knowledge of site reliability trends and shares new information with team members, management, and beyond to help others build, test, deploy and run services.
  • Performs standard and non‑standard analyses and provides clear data on production to contribute to business development decisions (e.g., design changes).
Core Responsibilities
  • Independently manages work, monitoring timelines and deliverables to ensure projects or initiatives stay on track and meet requirements. Proactively prioritizes work and adapts to resource or timeline shifts, suggesting adjustments to maintain project efficiency.
  • Collaborates across teams to align on expectations and achieve shared objectives. Builds and maintains a comprehensive understanding of business, stakeholder, and/or customer needs to build and support effective partnerships. Actively listens to diverse perspectives and asks questions to ensure understanding of others.
  • Independently identifies and addresses standard and non‑standard issues in accordance with standard practices, escalating more complex issues as appropriate. Analyzes data and/or information from multiple sources to troubleshoot standard and non‑standard errors. Contributes to knowledge sharing and best practices.
  • Embraces continuous learning by actively seeking to build knowledge and new skills and/or tools and staying current with industry trends and best practices. Seeks out and leverages feedback and training to improve skills. Contributes to a culture of continuous learning and knowledge sharing with team members.
  • Develops ideas and recommends updates to increase the efficiency and effectiveness of processes, protocols, and workflows within a team. Seeks input from team members on alternative approaches and methods for improving work.
Minimum Job Qualifications
  • 8 years of experience in software engineering, infrastructure management, or related field
  • Bachelor's Degree in Computer Science, Engineering, or related field AND 4 years of experience in software engineering, infrastructure management, or related field
  • Master's Degree in Computer Science, Engineering, or related field AND 2 year of experience in software engineering, infrastructure management, or related field.
  • Doctorate in Computer Science, Engineering, or related field.
Job Skills
  • Same skills as prior level plus;
  • Operating Systems Demonstrated ability in or knowledge of operating systems, including installing, upgrading, and troubleshooting various operating environments.
Automation Experience
  • 3 years of experience in automation.
Programming Experience
  • 3 years of experience in programming and/or scripting.
Preferred Job Qualifications
  • 9 years of experience in software engineering, infrastructure management, or related field
  • Bachelor's Degree in Computer Science, Engineering, or related field AND 5 years of experience in software engineering, infrastructure management, or related field
  • Master's Degree in Computer Science, Engineering, or related field AND 3 years of experience in software engineering, infrastructure management, or related field
  • Doctorate in Computer Science, Engineering, or related field AND 1 year of experience in software engineering, infrastructure management, or related field.
Automation Experience
  • 5 years of experience in automation.
Programming Experience
  • 5 years of experience in programming and/or scripting.
Qualifications

Disclaimer:

Certain U.S. based or U.S. customer or client‑facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.

Range and benefit information provided in this posting are specific to the stated locations only

US: Hiring Range in USD from: $81,100 to $187,000 per annum. May be eligible for bonus and equity.

Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.

Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:

  • Medical, dental, and vision insurance, including expert medical opinion
  • Short term disability and long term disability
  • Life insurance and AD&D
  • Supplemental life insurance (Employee/Spouse/Child)
  • Health care and dependent care Flexible Spending Accounts
  • Pre‑tax commuter and parking benefits
  • 401(k) Savings and Investment Plan with company match
  • Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non‑overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
  • 11 paid holidays
  • Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximium cap of 112 hours.
  • Paid parental leave
  • Adoption assistance
  • Employee Stock Purchase Plan
  • Financial planning and group legal
  • Voluntary benefits including auto homeowner and pet insurance

The role will generally accept applications for at least three calen