1

Prometheus Jobs in Atlanta, GA (NOW HIRING)

Openshift Devops Engineer

Atlanta, GA · On-site

$100K - $130K/yr

Experience with monitoring and logging tools such as Prometheus, Grafana, Splunk and Datadog. * Strong production troubleshooting skills across infrastructure, OS, containers, and applications.

Sr java Developer

Alpharetta, GA · On-site

$56 - $71.25/hr

Implement and enhance observability, including monitoring, alerting, and logging (e.g., Prometheus, Grafana, ELK, OpenTelemetry). Drive root cause analysis and performance tuning across services and ...

Senior Platform Engineer

East Point, GA · On-site

$101K - $169K/yr

Build and maintain the full observability stack: structured logging, Prometheus metrics, Kinesis Firehose OpenSearch indexing, and Grafana dashboards for per-user, per-tool, per-session audit trails

Senior Platform Engineer

Forest Park, GA · On-site

$101K - $169K/yr

Build and maintain the full observability stack: structured logging, Prometheus metrics, Kinesis Firehose OpenSearch indexing, and Grafana dashboards for per-user, per-tool, per-session audit trails

Senior Platform Engineer

Austell, GA · On-site

$101K - $169K/yr

Build and maintain the full observability stack: structured logging, Prometheus metrics, Kinesis Firehose OpenSearch indexing, and Grafana dashboards for per-user, per-tool, per-session audit trails

Senior Platform Engineer

Vinnings, GA · On-site

$101K - $169K/yr

Build and maintain the full observability stack: structured logging, Prometheus metrics, Kinesis Firehose OpenSearch indexing, and Grafana dashboards for per-user, per-tool, per-session audit trails

Showing results 21-40

Prometheus information

See Atlanta, GA salary details

$12

$20

$60

How much do prometheus jobs pay per hour?

As of Aug 22, 2026, the average hourly pay for prometheus in Atlanta, GA is $20.24, according to ZipRecruiter salary data. Most workers in this role earn between $13.85 and $16.63 per hour, depending on experience, location, and employer.

What is a Prometheus engineer?

Prometheus engineers are IT professionals who specialize in deploying, configuring, and maintaining Prometheus, an open-source systems monitoring and alerting toolkit. They design and implement monitoring solutions for infrastructure, applications, and services using Prometheus and its ecosystem (like Grafana for visualization). Their responsibilities include setting up metrics collection, writing queries, tuning alerts, and ensuring high availability and scalability of monitoring setups. Prometheus engineers often work closely with DevOps teams to provide observability and actionable insights for system reliability.

What skills and qualifications are needed to thrive as a Prometheus engineer?

To thrive as a Prometheus Monitoring Engineer, you need a solid understanding of systems administration, networking, and monitoring concepts, often supported by experience in DevOps or Site Reliability Engineering roles. Proficiency in Prometheus, Grafana, alerting tools, and infrastructure-as-code platforms, as well as familiarity with cloud services and container orchestration systems like Kubernetes, is essential. Strong problem-solving skills, attention to detail, and effective communication help you proactively address issues and collaborate with diverse technical teams. These skills ensure the reliability, scalability, and optimal performance of critical IT infrastructure.

What are common challenges faced by Prometheus engineers, and how can they be addressed?

Prometheus monitoring engineers often face challenges related to scaling the system to handle large volumes of metrics data and optimizing query performance. Dealing with high cardinality metrics and managing storage efficiently can also be complex. These challenges can be addressed by carefully designing label schemas, employing federation or sharding, and using long-term storage integrations. Collaboration with development and operations teams is key to ensure effective instrumentation and alerting, which helps maintain system reliability and performance.

What is the difference between Prometheus vs Grafana Developer?

AspectPrometheusGrafana Developer
Primary RoleMonitoring and alerting system configurationDashboard and visualization development
Required SkillsPromQL, server setup, metrics collectionGrafana, data sources, visualization design
Work EnvironmentDevOps, monitoring infrastructureData visualization, front-end development
CertificationsNone specific, familiarity with monitoring toolsGrafana certifications beneficial

Prometheus and Grafana Developer roles often overlap in monitoring and visualization. Prometheus focuses on metrics collection and alerting, while Grafana Developers specialize in creating dashboards for data visualization. Both roles are essential in DevOps environments but serve different functions within the monitoring ecosystem.

What are popular job titles related to Prometheus jobs in Atlanta, GA?

For Prometheus jobs in Atlanta, GA, the most frequently searched job titles are:

What job categories do people searching Prometheus jobs in Atlanta, GA look for?

The top searched job categories for Prometheus jobs in Atlanta, GA are:

Infographic showing various Prometheus job openings in Atlanta, GA as of August 2026, with employment types broken down into 90% Full Time, 3% Part Time, and 7% Contract. Highlights an 80% Physical, 6% Hybrid, and 14% Remote job distribution, with an average salary of $42,099 per year, or $20.2 per hour.

Site Reliability Engineer (SRE) - AI Platform & Cloud

Morgan Stanley

Alpharetta, GA • On-site

$54.25 - $72/hr

Full-time

Re-posted 2 days ago


Morgan Stanley rating

8.4

Company rating: 8.4 out of 10

Based on 155 frontline employees who took The Breakroom Quiz

33rd of 151 rated financial services


Job description

Job Summary:
Morgan Stanley is a global leader in financial services, known for its innovative approach to technology. They are seeking a Director-level Site Reliability Engineer (SRE) to join their AI Platform team, responsible for maintaining and scaling the infrastructure that supports AI/ML systems in a high-stakes financial environment.
Responsibilities:
• Operate, monitor, and maintain the infrastructure supporting GenAI applications (training, inference, feature store, data ingestion, model serving)
• Design and build automation for core platform capabilities, reducing manual toil
• Develop and maintain infrastructure-as-code (IaC) for provisioning and managing compute, storage, network, GPU clusters, Kubernetes / container orchestration, etc.
• Establish, monitor, and enforce SLOs/SLIs/SLAs, error budgets, alerting, and dashboards
• Lead incident response, root cause analysis (RCA), postmortems, and systemic remediation
• Perform capacity planning, scaling strategies, workload scheduling, and resource forecasting
• Optimize cost vs. performance tradeoffs in large-scale compute environments
• Harden systems for security, compliance, auditability, and data governance
• Collaborate across teams (cloud engineers, data engineers, infrastructure, security) to ensure safe deployment, rollout, rollback, and integration of new systems
• Define disaster recovery (DR) strategies, backup/restore practices, fault tolerance mechanisms
• Maintain runbooks, operational playbooks, documentation, and training materials
• Participate in on-call rotations and respond to production incidents 24/7 as needed
• Continuously evaluate and integrate new tools, frameworks, or technologies to enhance platform reliability
Qualifications:
Required:
• Bachelor’s or Master’s degree in Computer Science or related field, or equivalent job experience
• 5 years of production experience in SRE / Infrastructure / ops for large-scale systems
• Strong programming/scripting skills (Python, Go, Java, or equivalent)
• Deep experience with containerization (Docker), orchestration (Kubernetes, etc.)
• Infrastructure-as-code (Terraform, Helm, CloudFormation, Ansible, etc.)
• Familiarity with GPU / AI compute clusters, high-performance data storage, and distributed architectures
• Experience with monitoring / observability / logging / alerting tools (Prometheus, Grafana, ELK / EFK, Datadog, etc.)
• Networking & systems engineering knowledge (TCP/IP, DNS, routing, load balancing, distributed storage)
• Solid experience in capacity planning, performance tuning, scaling, and incident response
• Demonstrated ability to lead RCAs, deploy fixes, and drive reliability improvements
• Experience in regulated environments (financial services, compliance, audit, security) is a strong plus
• Excellent communication, documentation, and cross-team collaboration skills
• Proven track record of reducing operational toil via automation
Preferred:
• Understanding of SRE techniques.
• Proficiency with Open Telemetry tools including Grafana, Loki, Prometheus, and Cortex.
• Good knowledge of Microservice based architecture, industry standards, for both public and private cloud.
• Knowledge of data pipeline technologies (Kafka, Spark, Flink, etc.)
• Good knowledge of various DB engines (SQL, Redis, Kafka, Snowflake, etc) for cloud app storage.
• Experience working with Generative AI development, embeddings, fine tuning of Generative AI models.
• Experience in high-performance computing (HPC), distributed GPU cluster scheduling (e.g. Slurm, Kubernetes GPU scheduling)
• Understanding of ModelOps/ ML Ops/ LLM Op.
• Experience with chaos engineering, canary deployments, blue/green rollouts
Company:
Morgan Stanley is a financial services institution that delivers capital management, investment banking, and advisory solutions. Founded in 1935, the company is headquartered in New York, USA, with a team of 10001+ employees. The company is currently Late Stage.

What Morgan Stanley employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Morgan Stanley logo

About Morgan Stanley

Sourced by ZipRecruiter

Since our founding in 1935, Morgan Stanley has been committed to serving local and global communities by being a market leader in Investment Banking, Securities, Investment Management and Wealth Management services. Our belief that capital can work to benefit all of society inspires us to put our clients first, lead with exceptional ideas, hold our business to high ethical standards, and give back to communities around the world through philanthropy and public works. We have a smart casual dress code and operate under a philosophy that balances work with your personal life. Our people's talent, passion, and expertise is the fuel on which our organization runs, therefore, our people are our greatest asset. Diversity and inclusiveness is a critical component for our success and it is our priority to continue building a firm that values the unique background and identity of every one of our employees, thus enabling our people to bring their full, and best selves to work each day. Teamwork is the essence of our approach, and so are the values of integrity, excellence, and enabling our people to achieve at the highest levels. We invite you to learn more about our commitment to diversity and serving our community.

Industry

Finance and insurance and software development

Company size

10,000+ Employees

Headquarters location

New York, NY, US