1

Process Reliability Manager Jobs in Buffalo, NY (NOW HIRING)

Site Reliability Engineer

Buffalo, NY · On-site

$55.25 - $73.50/hr

Experience working with source code management tools and deployment processes. * Experience ... Experience with Site Reliability Engineering practices, including SLOs, SLIs, SLAs, error budgets ...

Senior OES Process Engineer

Tonawanda, NY

$99K - $128K/yr

... safety, reliability, efficiency, and performance. * Provide remote and on-site support to ... Collaborate with operations technicians, plant management, regional operations teams, logistics ...

Senior OES Process Engineer

Tonawanda, NY · On-site

$99K - $128K/yr

... safety, reliability, efficiency, and performance. * Provide remote and on-site support to ... Collaborate with operations technicians, plant management, regional operations teams, logistics ...

Managing documentation to include work instructions, process standards, technical documents, etc ... reliability engineer in a private, public, government or military environment. Additional ...

Managing documentation to include work instructions, process standards, technical documents, etc ... reliability engineer in a private, public, government or military environment. Additional ...

Senior OES Process Engineer

Tonawanda, NY · On-site

$108K - $158K/yr

... safety, reliability, efficiency, and performance. * Provide remote and on-site support to ... Collaborate with operations technicians, plant management, regional operations teams, logistics ...

next page

Showing results 1-20

Process Reliability Manager information

See Buffalo, NY salary details

$60.1K

$113.8K

$163.2K

How much do process reliability manager jobs pay per year?

As of Jul 28, 2026, the average yearly pay for process reliability manager in Buffalo, NY is $113,806.00, according to ZipRecruiter salary data. Most workers in this role earn between $91,500.00 and $135,600.00 per year, depending on experience, location, and employer.

What is a Process Reliability Manager?

A Process Reliability Manager is a professional responsible for ensuring that manufacturing or production processes operate efficiently, consistently, and with minimal downtime. They analyze process data, identify areas for improvement, and implement strategies to enhance equipment reliability and overall process performance. By collaborating with maintenance, engineering, and operations teams, they help reduce failures, optimize productivity, and maintain quality standards. Their work is crucial for minimizing costs and ensuring that production targets are met safely and reliably.

What is the difference between Process Reliability Manager vs Maintenance Engineer?

AspectProcess Reliability ManagerMaintenance Engineer
CertificationsReliability certifications, Six Sigma, PMPMechanical/Electrical certifications, HVAC, PLC certifications
Work EnvironmentManufacturing plants, industrial facilitiesFactories, equipment maintenance sites
Industry UsageFocus on reliability, uptime, and process optimizationFocus on equipment repair, preventive maintenance

The Process Reliability Manager primarily focuses on improving equipment reliability and process efficiency through data analysis and strategic planning. In contrast, Maintenance Engineers handle the hands-on repair and maintenance of machinery. Both roles are essential in manufacturing, but the Process Reliability Manager emphasizes proactive reliability strategies, while Maintenance Engineers focus on reactive and preventive maintenance tasks.

How does a Process Reliability Manager typically collaborate with maintenance and production teams to achieve operational goals?

A Process Reliability Manager works closely with both maintenance and production teams to identify areas of improvement in equipment reliability and process efficiency. This often involves facilitating cross-functional meetings, analyzing downtime data, and implementing preventive maintenance strategies. Clear communication and teamwork are key, as the role requires aligning the objectives of different departments to minimize unplanned outages and optimize production output. By fostering a proactive culture and sharing best practices, the Process Reliability Manager helps ensure the plant operates smoothly and efficiently.

What are the key skills and qualifications needed to thrive as a Process Reliability Manager, and why are they important?

To thrive as a Process Reliability Manager, you need a strong background in engineering, process optimization, and reliability analysis, often supported by a degree in engineering and experience in manufacturing or industrial settings. Familiarity with reliability-centered maintenance (RCM), root cause analysis tools, and data analysis software such as SAP or Maximo is typically required. Exceptional problem-solving, leadership, and communication skills help drive cross-functional teams and foster a culture of continuous improvement. These skills are crucial to ensure equipment reliability, minimize downtime, and optimize operational efficiency within complex production environments.
What cities near Buffalo, NY are hiring for Process Reliability Manager jobs? Cities near Buffalo, NY with the most Process Reliability Manager job openings:
Infographic showing various Process Reliability Manager job openings in Buffalo, NY as of July 2026, with employment types broken down into 83% Full Time, 15% Part Time, 1% Temporary, and 1% Contract. Highlights an 86% Physical, 1% Hybrid, and 13% Remote job distribution, with an average salary of $113,806 per year, or $54.7 per hour.
Solution Architect AI Platform Reliability & SRE (Mythos SRE)

Solution Architect AI Platform Reliability & SRE (Mythos SRE)

Imagine Staffing Technology

Buffalo, NY • Remote

$55.25 - $73.50/hr

Full-time

Posted 28 days ago


Job description

Job Title: Solution Architect – AI Platform Reliability & SRE (Mythos SRE)
Location: Remote (Within USA)
Hire Type: Contract
Pay Range: Competitive Hourly Rate
Work Model: Remote with periodic travel to Buffalo, NY
Schedule: Monday – Friday, Standard Business Hours
Recruiter Contact: Samantha Marranca | 716-256-1271 | smarranca@imaginestaffing.net
NO C2C, NO sponsorship given at this time
Nature & Scope:
Positional Overview
Our client is seeking an experienced Solution Architect to support the reliability, scalability, observability, and operational excellence of its enterprise AI platform, Mythos. This role serves as the solution architecture extension of Enterprise Architecture and AI Platform teams, translating strategic platform designs into detailed operational architectures that enable highly available, resilient, and scalable AI services.
The Solution Architect will partner closely with Site Reliability Engineering (SRE), Platform Engineering, Infrastructure, Cloud Operations, and Application Development teams to establish architecture patterns and operational frameworks that support enterprise AI workloads across cloud and co-location environments.
This position is ideal for a hands-on architect with expertise in cloud infrastructure, platform engineering, observability, reliability engineering, and large-scale distributed systems.
Role & Responsibility:
Tasks That Will Lead To Your Success
AI Platform Reliability Architecture
  • Translate enterprise AI platform architecture into detailed operational and infrastructure solution designs.
  • Define reliability, scalability, resiliency, and availability architecture standards for AI workloads.
  • Develop architecture patterns supporting highly available and fault-tolerant AI services.
  • Support enterprise AI platform growth through scalable infrastructure and platform design.
  • Establish architecture guidance for production readiness and operational excellence.
Site Reliability Engineering & Operational Excellence
  • Define architecture patterns supporting SRE best practices across AI platforms.
  • Support implementation of Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budget frameworks.
  • Develop operational readiness standards and deployment validation processes.
  • Establish reliability engineering practices that improve system stability and performance.
  • Partner with engineering teams to improve incident prevention, detection, and response capabilities.
Scalability & Performance Optimization
  • Design solutions supporting large-scale AI workloads and model-serving environments.
  • Establish architecture patterns that optimize platform performance and resource utilization.
  • Support capacity planning and infrastructure scaling strategies.
  • Identify performance bottlenecks and recommend architectural improvements.
  • Collaborate with engineering teams to improve application and platform efficiency.
Observability & Monitoring
  • Design enterprise observability frameworks supporting AI platform operations.
  • Establish telemetry standards providing visibility into system health, model performance, operational metrics, and risk indicators.
  • Define monitoring, alerting, logging, and tracing strategies.
  • Support implementation of observability tools and telemetry platforms.
  • Ensure operational teams have actionable insights supporting platform reliability and performance.
Infrastructure & Automation
  • Develop architecture guidance for Infrastructure as Code (IaC) and platform automation.
  • Support CI/CD pipeline architecture and deployment automation strategies.
  • Establish repeatable operational patterns supporting cloud and co-location environments.
  • Promote infrastructure standardization and operational consistency.
  • Collaborate with Platform Engineering teams on automation and operational tooling initiatives.
AI Operational Governance
  • Support architecture strategies for AI model monitoring and drift detection.
  • Establish operational frameworks supporting AI governance and platform controls.
  • Define reliability patterns for embedded AI capabilities within enterprise applications.
  • Ensure platform operations align with enterprise security, compliance, and risk management standards.
Cross-Functional Collaboration
  • Partner with Enterprise Architects, Platform Engineering, Infrastructure, Security, Observability, and Development teams.
  • Participate in architecture reviews, design workshops, and Agile ceremonies.
  • Provide technical guidance throughout the SDLC from design through production deployment.
  • Validate architecture decisions and ensure adherence to enterprise reliability standards.
  • Contribute operational insights that influence future platform architecture decisions.
Skills & Experience
Qualifications That Will Help You Thrive
Required Experience
  • Bachelor’s Degree in Computer Science, Information Technology, Engineering, or related discipline.
  • 5+ years of experience in Solution Architecture, Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Architecture.
  • Experience designing highly available, scalable, and resilient distributed systems.
  • Strong understanding of cloud infrastructure and platform architecture principles.
  • Experience supporting production operations and enterprise-scale technology environments.
  • Knowledge of observability, monitoring, logging, and telemetry frameworks.
  • Experience with Infrastructure as Code and deployment automation concepts.
  • Strong communication and stakeholder management skills.
Preferred Qualifications
Experience supporting AI, Machine Learning, or Generative AI platforms.
Experience with Kubernetes, container orchestration, and cloud-native technologies.
Familiarity with observability platforms such as Datadog, Dynatrace, Grafana, Prometheus, Splunk, or OpenTelemetry.
Experience implementing SLI, SLO, and error budget frameworks.
Experience with Infrastructure as Code technologies such as Terraform or CloudFormation.
Cloud certifications within Azure, AWS, or Google Cloud.
Financial services experience preferred.
Experience supporting highly regulated enterprise environments.
Team & Environment
Works closely with Enterprise Architecture, Platform Engineering, Infrastructure, SRE, Security, and Application Development teams.
Serves as a key architecture resource supporting enterprise AI platform operations.
Participates in highly collaborative Agile teams.
Provides technical leadership supporting reliability and operational excellence initiatives.
Work Schedule & Travel
Schedule
  • Monday – Friday
  • Standard business hours
  • Flexible remote work environment
Travel
  • Occasional travel to Buffalo, NY
  • Approximately every 4–6 weeks as required
Compensation & Benefits
  • Competitive hourly compensation
  • Long-term contract engagement
  • Remote work flexibility
  • Opportunity to influence enterprise-wide AI strategy and adoption
Why Join This Opportunity?
This is a unique opportunity to help build and operate next-generation AI platforms at enterprise scale. The successful candidate will play a critical role in ensuring the reliability, resilience, observability, and operational success of AI technologies that support strategic business initiatives across the organization.