1

Infrastructure Monitoring Engineer Jobs in Tennessee

Infrastructure Engineer

Nashville, TN · On-site

$103K - $136K/yr

Monitoring these operating environments. * Responding effectively and speedily to any problems ... Infrastructure Engineer Requirements: * Bachelor's degree in computer science or engineering.

Infrastructure Platform Engineer

Oak Ridge, TN · On-site +1

$102K - $134K/yr

... an HPC Infrastructure Platform Engineer to join the HPC Infrastructure group. The preferred ... Monitor systems using tools like Nagios and Grafana * Respond to and assist in troubleshooting ...

HPC Infrastructure Platform Engineer

Knoxville, TN · On-site

$95K - $125K/yr

HPC Infrastructure Platform Engineer Founded in 1999 in the beautiful Smoky Mountains of East ... Monitor systems using tools like Nagios and Grafana * Respond to and assist in troubleshooting ...

... an HPC Infrastructure Platform Engineer to join the HPC Infrastructure group. The preferred ... Monitor systems using tools like Nagios and Grafana * Respond to and assist in troubleshooting ...

next page

Showing results 1-20

Infrastructure Monitoring Engineer information

What does an infrastructure monitoring engineer do?

An Infrastructure Monitoring Engineer is responsible for overseeing the performance and availability of an organization's IT systems, networks, servers, and applications. They use specialized monitoring tools to detect and respond to issues, ensuring minimal downtime and optimal system functioning. These engineers analyze data, set up alerts, and work with other IT teams to resolve problems quickly. Their role is crucial for maintaining the reliability and efficiency of complex infrastructures.

What are some common challenges infrastructure monitoring engineers face when supporting large-scale systems?

Infrastructure Monitoring Engineers often encounter challenges such as managing the sheer volume of alerts generated by monitoring tools, ensuring coverage across diverse and evolving technology stacks, and maintaining high system availability. Balancing proactive detection of potential issues with minimizing false positives is crucial to avoid alert fatigue. Additionally, effective collaboration with IT, operations, and development teams is essential for swift incident resolution and continuous improvement of monitoring strategies.

What are the key skills and qualifications needed to thrive as an infrastructure monitoring engineer, and why are they important?

To thrive as an Infrastructure Monitoring Engineer, you need a strong background in IT systems, networking, and troubleshooting, often supported by relevant certifications like CompTIA Network+ or Microsoft Certified: Azure Administrator. Familiarity with monitoring tools such as Nagios, Zabbix, or Datadog, as well as scripting languages like Python or PowerShell, is typically required. Strong analytical thinking, attention to detail, and effective communication set top professionals apart in this field. These skills and qualities are critical for proactively identifying and resolving issues to ensure optimal system performance and minimize downtime.

What is the difference between Infrastructure Monitoring Engineer vs Network Operations Center (NOC) Technician?

AspectInfrastructure Monitoring EngineerNetwork Operations Center (NOC) Technician
CertificationsCompTIA Network+, Cisco CCNA, monitoring tools certificationsCompTIA Network+, Cisco CCNA, network monitoring certifications
Work EnvironmentData centers, cloud environments, enterprise IT infrastructureNetwork operation centers, service provider facilities
Industry UsageIT, cloud services, enterprise infrastructureTelecommunications, internet service providers, enterprise networks
Primary FocusMonitoring infrastructure health, performance, and availabilityMonitoring network traffic, troubleshooting connectivity issues

The Infrastructure Monitoring Engineer focuses on overseeing the health and performance of IT infrastructure components, including servers and cloud systems. In contrast, the NOC Technician primarily monitors network traffic and connectivity issues within network operations centers. While both roles require similar certifications and work environments, their core responsibilities differ in scope and focus.

What are popular job titles related to Infrastructure Monitoring Engineer jobs in Tennessee?

For Infrastructure Monitoring Engineer jobs in Tennessee, the most frequently searched job titles are:

What job categories do people searching Infrastructure Monitoring Engineer jobs in Tennessee look for?

The top searched job categories for Infrastructure Monitoring Engineer jobs in Tennessee are:

What cities in Tennessee are hiring for Infrastructure Monitoring Engineer jobs?

Cities in Tennessee with the most Infrastructure Monitoring Engineer job openings:

Infographic showing various Infrastructure Monitoring Engineer job openings in Tennessee as of June 2026, with employment types broken down into 83% Full Time, 15% Part Time, and 2% Contract. Highlights an 88% Physical, 3% Hybrid, and 9% Remote job distribution.

Systems Development Engineer , Operations Infrastructure Services

Socket.dev

Nashville, TN • On-site

$123 - $166/hr

Other

Medical, Dental, Vision, Life, Retirement, PTO

Posted 5 days ago


Key responsibilities

  • Design, build, and operate scalable monitoring infrastructure on AWS that ingests and processes high-volume device telemetry and network topology data

  • Own systems end‑to‑end across monitoring tooling and data pipelines, from build through deployment, operation, and ongoing maintenance

  • Partner with infrastructure and operations stakeholders to grasp monitoring gaps and translate them into reliable, automated solutions that reduce manual intervention


Job description

Join us in building infrastructure monitoring applications that power Amazon's global operations. You'll design and deliver services that process high-volume device telemetry, evaluate network linkages in real time, and provide operators with clear, actionable insights across thousands of sites worldwide.

We're an agile development team within Operations Infrastructure Services (OIS), part of Amazon Robotics, assisting fulfillment centers, delivery stations, and sortation centers globally. Our monitoring product gives operators a single, reliable view of device and network health—from live infrastructure telemetry to reliance‑aware alarming that points to root causes instead of overwhelming teams with duplicate alerts. You'll have the opportunity to apply modern AI technologies, from AI‑assisted incident detection to generative AI tooling that accelerates how we build and operate our network.

You might start your day reviewing a design for integrating metrics from a new device type being onboarded across OIS sites worldwide, then inspect runtime metrics to tune collection thresholds and data quality before shipping a monitoring improvement operators experience immediately. You'll partner with teammates on design reviews, participate in quick feedback loops, and own features end‑to‑end — working alongside infrastructure teams to deliver the monitoring experience needed to reduce building downtime and assistance. Throughout the day, you'll balance designing scalable monitoring solutions with hands‑on implementation, working across back‑end services and data pipelines while assisting each other's growth and bringing your authentic perspective to the team.

Key job responsibilities
  • Design, build, and operate scalable monitoring infrastructure on AWS that ingests and processes high-volume device telemetry and network topology data
  • Own systems end‑to‑end across monitoring tooling and data pipelines, from build through deployment, operation, and ongoing maintenance
  • Partner with infrastructure and operations stakeholders to grasp monitoring gaps and translate them into reliable, automated solutions that reduce manual intervention
  • Apply AI/ML and generative AI techniques to improve detection quality, reduce alarm noise, and streamline operational workflows
  • Raise the bar on system reliability, operational excellence and automation through thoughtful design, thorough testing, and continuous improvement
A day in the life

You might start by reviewing a design for integrating metrics from a new device type being onboarded across all Robotics buildings worldwide, then inspect runtime metrics to tune metric collection and quality before shipping a service improvement our operators experience immediately. Our team values partnership, quick feedback loops, and clear ownership.

Amazon offers a full range of benefits that assist you and eligible family members, including domestic partners. Benefits can vary by location, the number of regularly scheduled hours you work, length of employment, and job status such as seasonal or temporary employment. The benefits that generally apply to regular, full‑time employees include:

  1. Medical, Dental, and Vision Coverage
  2. Maternity and Parental Leave Options
  3. Paid Time Off (PTO)
  4. 401(k) Plan

If you are not sure that every qualification on the list above describes you exactly, we'd still love to hear from you! At Amazon, we value people with unique backgrounds, experiences, and skillsets.

About the team

We're a close‑knit, agile team that owns infrastructure monitoring for OIS within Amazon Robotics and ships to a worldwide operational fleet. We care deeply about the craft of software and foster each other's growth. Our vision centers on delivering reliable, intelligent monitoring that helps operators grasp their infrastructure at scale. You'll work with product and operations stakeholders to turn complex monitoring challenges into simple, elegant solutions. We're excited to welcome someone who aligns with our standards for well‑architected software and collective problem‑solving.

Basic Qualifications
  • 3+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience
  • Experience in automating, deploying, and supporting large‑scale infrastructure
  • Experience programming with at least one modern language such as Python, Ruby, Golang, Java, C++, C#, Rust
  • Experience with Linux/Unix
  • Experience with CI/CD pipelines build processes
Preferred Qualifications
  • 3+ years of non‑internship professional software development experience
  • Experience with distributed systems at scale

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign‑on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, mental health support, medical advice line, flexible spending accounts, adoption and surrogacy reimbursement coverage), 401(k) matching, paid time off, and parental leave.

Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, TN, Nashville - 122,800.00 - 166,100.00 USD annually

#J-18808-Ljbffr