1

Manager Hpc Engineer Jobs in Georgia (NOW HIRING)

Service Engineer

Atlanta, GA · On-site

$70K - $100K/yr

... HPC and IoT/Embedded customers worldwide. We are the #5 fastest growing company among the Silicon ... managing repair/parts cycle times • Ensure escalation situations are managed and corrected ...

Oversee process integration, manage scheduling, and provide comprehensive progress reports to upper ... Familiarity with bifurcated roadmap considerations (high-performance AI/HPC vs. cost-optimized ...

Oversee process integration, manage scheduling, and provide comprehensive progress reports to upper ... Familiarity with bifurcated roadmap considerations (high-performance AI/HPC vs. cost-optimized ...

Oversee process integration, manage scheduling, and provide comprehensive progress reports to upper ... Familiarity with bifurcated roadmap considerations (high-performance AI/HPC vs. cost-optimized ...

Senior AI Engineer - SFL Scientific

Atlanta, GA · On-site

$100K - $138K/yr

Ability to manage and prioritize multiple tasks in a fast-paced and dynamic environment * Strong ... and HPC system software stack The wage range for this role takes into account the wide range of ...

Adopt best engineering practices in automation, HPC and AI/GenAI infrastructure and design patterns ... Ability to manage and prioritize multiple tasks in a fast-paced and dynamic environment * Strong ...

... HPC), artificial intelligence (AI), machine learning (ML), and advanced data analytics to design ... Youwillhelp users configure, manage, andoptimizetheiruse of custom APIs,anddatasets.You will design ...

Posted today

Sr Sales Engineer - US-East

Atlanta, GA · On-site

$175.75 - $219.69/hr

Support lead generation and pipeline development alongside territory managers, including trade show ... Familiarity with HPC, scientific computing, or performance engineering use cases, including ...

New

Showing results 21-40

Manager Hpc Engineer information

What is the difference between Manager Hpc Engineer vs Hpc Engineer?

AspectManager Hpc EngineerHpc Engineer
CredentialsBachelor's/Master's in Computer Science or related, often with leadership experienceBachelor's or higher in Computer Science, Engineering, or related
Work EnvironmentLeads teams, manages projects, oversees HPC system deployment and maintenanceDesigns, develops, and maintains HPC systems and applications
Employer & Industry UsageUsed in research institutions, tech companies, and data centers with a focus on team managementCommon in scientific research, academia, and enterprise sectors focusing on HPC infrastructure

The main difference between a Manager Hpc Engineer and an Hpc Engineer is that the manager oversees teams and projects, focusing on leadership and strategic planning, while the Hpc Engineer concentrates on technical design, implementation, and maintenance of HPC systems. Both roles require strong technical skills, but the manager also needs leadership and project management abilities.

What are the most commonly searched types of Hpc Engineer jobs in Georgia?

The most popular types of Hpc Engineer jobs in Georgia are:

What cities in Georgia are hiring for Manager Hpc Engineer jobs?

Cities in Georgia with the most Manager Hpc Engineer job openings:

Site Reliability Engineer (SRE) - AI Platform & Cloud

Morgan Stanley

Alpharetta, GA • On-site

$54.25 - $72/hr

Full-time

Re-posted yesterday


Morgan Stanley rating

8.4

Company rating: 8.4 out of 10

Based on 155 frontline employees who took The Breakroom Quiz

33rd of 151 rated financial services


Job description

Job Summary:
Morgan Stanley is a global leader in financial services, known for its innovative approach to technology. They are seeking a Director-level Site Reliability Engineer (SRE) to join their AI Platform team, responsible for maintaining and scaling the infrastructure that supports AI/ML systems in a high-stakes financial environment.
Responsibilities:
• Operate, monitor, and maintain the infrastructure supporting GenAI applications (training, inference, feature store, data ingestion, model serving)
• Design and build automation for core platform capabilities, reducing manual toil
• Develop and maintain infrastructure-as-code (IaC) for provisioning and managing compute, storage, network, GPU clusters, Kubernetes / container orchestration, etc.
• Establish, monitor, and enforce SLOs/SLIs/SLAs, error budgets, alerting, and dashboards
• Lead incident response, root cause analysis (RCA), postmortems, and systemic remediation
• Perform capacity planning, scaling strategies, workload scheduling, and resource forecasting
• Optimize cost vs. performance tradeoffs in large-scale compute environments
• Harden systems for security, compliance, auditability, and data governance
• Collaborate across teams (cloud engineers, data engineers, infrastructure, security) to ensure safe deployment, rollout, rollback, and integration of new systems
• Define disaster recovery (DR) strategies, backup/restore practices, fault tolerance mechanisms
• Maintain runbooks, operational playbooks, documentation, and training materials
• Participate in on-call rotations and respond to production incidents 24/7 as needed
• Continuously evaluate and integrate new tools, frameworks, or technologies to enhance platform reliability
Qualifications:
Required:
• Bachelor’s or Master’s degree in Computer Science or related field, or equivalent job experience
• 5 years of production experience in SRE / Infrastructure / ops for large-scale systems
• Strong programming/scripting skills (Python, Go, Java, or equivalent)
• Deep experience with containerization (Docker), orchestration (Kubernetes, etc.)
• Infrastructure-as-code (Terraform, Helm, CloudFormation, Ansible, etc.)
• Familiarity with GPU / AI compute clusters, high-performance data storage, and distributed architectures
• Experience with monitoring / observability / logging / alerting tools (Prometheus, Grafana, ELK / EFK, Datadog, etc.)
• Networking & systems engineering knowledge (TCP/IP, DNS, routing, load balancing, distributed storage)
• Solid experience in capacity planning, performance tuning, scaling, and incident response
• Demonstrated ability to lead RCAs, deploy fixes, and drive reliability improvements
• Experience in regulated environments (financial services, compliance, audit, security) is a strong plus
• Excellent communication, documentation, and cross-team collaboration skills
• Proven track record of reducing operational toil via automation
Preferred:
• Understanding of SRE techniques.
• Proficiency with Open Telemetry tools including Grafana, Loki, Prometheus, and Cortex.
• Good knowledge of Microservice based architecture, industry standards, for both public and private cloud.
• Knowledge of data pipeline technologies (Kafka, Spark, Flink, etc.)
• Good knowledge of various DB engines (SQL, Redis, Kafka, Snowflake, etc) for cloud app storage.
• Experience working with Generative AI development, embeddings, fine tuning of Generative AI models.
• Experience in high-performance computing (HPC), distributed GPU cluster scheduling (e.g. Slurm, Kubernetes GPU scheduling)
• Understanding of ModelOps/ ML Ops/ LLM Op.
• Experience with chaos engineering, canary deployments, blue/green rollouts
Company:
Morgan Stanley is a financial services institution that delivers capital management, investment banking, and advisory solutions. Founded in 1935, the company is headquartered in New York, USA, with a team of 10001+ employees. The company is currently Late Stage.

What Morgan Stanley employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Morgan Stanley logo

About Morgan Stanley

Sourced by ZipRecruiter

Since our founding in 1935, Morgan Stanley has been committed to serving local and global communities by being a market leader in Investment Banking, Securities, Investment Management and Wealth Management services. Our belief that capital can work to benefit all of society inspires us to put our clients first, lead with exceptional ideas, hold our business to high ethical standards, and give back to communities around the world through philanthropy and public works. We have a smart casual dress code and operate under a philosophy that balances work with your personal life. Our people's talent, passion, and expertise is the fuel on which our organization runs, therefore, our people are our greatest asset. Diversity and inclusiveness is a critical component for our success and it is our priority to continue building a firm that values the unique background and identity of every one of our employees, thus enabling our people to bring their full, and best selves to work each day. Teamwork is the essence of our approach, and so are the values of integrity, excellence, and enabling our people to achieve at the highest levels. We invite you to learn more about our commitment to diversity and serving our community.

Industry

Finance and insurance and software development

Company size

10,000+ Employees

Headquarters location

New York, NY, US