Support reliability engineering practices to improve uptime, resilience, and service continuity ... Monitoring, observability, and incident response tooling. * Networking fundamentals: DNS, routing ...
Support reliability engineering practices to improve uptime, resilience, and service continuity ... Monitoring, observability, and incident response tooling. * Networking fundamentals: DNS, routing ...
Support reliability engineering practices to improve uptime, resilience, and service continuity ... Monitoring, observability, and incident response tooling. * Networking fundamentals: DNS, routing ...
Support reliability engineering practices to improve uptime, resilience, and service continuity ... Monitoring, observability, and incident response tooling. * Networking fundamentals: DNS, routing ...
Support reliability engineering practices to improve uptime, resilience, and service continuity ... Monitoring, observability, and incident response tooling. * Networking fundamentals: DNS, routing ...
Support reliability engineering practices to improve uptime, resilience, and service continuity ... Monitoring, observability, and incident response tooling. * Networking fundamentals: DNS, routing ...
Own the reliability, performance, cost, and observability of deployed APIs, and pipelines (logging, metrics, tracing, alerting, evals). * Drive engineering best practices across the team: code review ...
Own the reliability, performance, cost, and observability of deployed APIs, and pipelines (logging, metrics, tracing, alerting, evals). * Drive engineering best practices across the team: code review ...
Own the reliability, performance, cost, and observability of deployed APIs, and pipelines (logging, metrics, tracing, alerting, evals). * Drive engineering best practices across the team: code review ...
Own the reliability, performance, cost, and observability of deployed APIs, and pipelines (logging, metrics, tracing, alerting, evals). * Drive engineering best practices across the team: code review ...
Own the reliability, performance, cost, and observability of deployed APIs, and pipelines (logging, metrics, tracing, alerting, evals). * Drive engineering best practices across the team: code review ...
Own the reliability, performance, cost, and observability of deployed APIs, and pipelines (logging, metrics, tracing, alerting, evals). * Drive engineering best practices across the team: code review ...
Senior Advanced AI Research Engineer
Oklahoma City, OK · On-site
$97K - $133K/yr
... horizon task reliability. The specific focus areas evolve with client demand and research ... Instrument research prototypes with observability from the start - distributed tracing, cost ...
Senior Advanced AI Research Engineer
Oklahoma City, OK · On-site
$97K - $133K/yr
... horizon task reliability. The specific focus areas evolve with client demand and research ... Instrument research prototypes with observability from the start - distributed tracing, cost ...
The Electrical and Instrument Reliability Program Manager reports to the Vice President of ... for site engineers. This position will serve as a subject matter expert on reliable design ...
The Electrical and Instrument Reliability Program Manager reports to the Vice President of ... for site engineers. This position will serve as a subject matter expert on reliable design ...
$91K - $125K/yr
... horizon task reliability. The specific focus areas evolve with client demand and research ... Instrument research prototypes with observability from the start - distributed tracing, cost ...
$91K - $125K/yr
... horizon task reliability. The specific focus areas evolve with client demand and research ... Instrument research prototypes with observability from the start - distributed tracing, cost ...
The Electrical and Instrument Reliability Program Manager reports to the Vice President of ... for site engineers. This position will serve as a subject matter expert on reliable design ...
The Electrical and Instrument Reliability Program Manager reports to the Vice President of ... for site engineers. This position will serve as a subject matter expert on reliable design ...
... reliability, security, and developer velocity. You champion build-it / run-it ownership, ensuring services meet clear standards for quality, observability, and operational excellence. You accelerate ...
... reliability, security, and developer velocity. You champion build-it / run-it ownership, ensuring services meet clear standards for quality, observability, and operational excellence. You accelerate ...
... reliability, security, and developer velocity. You champion buildit / runit ownership, ensuring services meet clear standards for quality, observability, and operational excellence. You accelerate ...
... reliability, security, and developer velocity. You champion buildit / runit ownership, ensuring services meet clear standards for quality, observability, and operational excellence. You accelerate ...
... reliability, security, and developer velocity. You champion build‑it / run‑it ownership, ensuring services meet clear standards for quality, observability, and operational excellence. You ...
... reliability, security, and developer velocity. You champion build‑it / run‑it ownership, ensuring services meet clear standards for quality, observability, and operational excellence. You ...
This role troubleshoots equipment issues, analyzes reliability data, supports maintenance ... Support site rotating equipment, including pumps, compressors, blowers, turbines, gearboxes, seals ...
This role troubleshoots equipment issues, analyzes reliability data, supports maintenance ... Support site rotating equipment, including pumps, compressors, blowers, turbines, gearboxes, seals ...
This role troubleshoots equipment issues, analyzes reliability data, supports maintenance ... Support site rotating equipment, including pumps, compressors, blowers, turbines, gearboxes, seals ...
This role troubleshoots equipment issues, analyzes reliability data, supports maintenance ... Support site rotating equipment, including pumps, compressors, blowers, turbines, gearboxes, seals ...
Rotating Equipment Engineer
Claremore, OK · On-site
This role troubleshoots equipment issues, analyzes reliability data, supports maintenance ... Support site rotating equipment, including pumps, compressors, blowers, turbines, gearboxes, seals ...
Rotating Equipment Engineer
Claremore, OK · On-site
This role troubleshoots equipment issues, analyzes reliability data, supports maintenance ... Support site rotating equipment, including pumps, compressors, blowers, turbines, gearboxes, seals ...
Machine Learning Engineer (AWS)
Tulsa, OK · On-site
Continuously explore new AWS services, ML frameworks, and deployment patterns to improve reliability, observability, and developer velocity on the ML platform. Requirements * 3+ years of experience ...
Machine Learning Engineer (AWS)
Tulsa, OK · On-site
Continuously explore new AWS services, ML frameworks, and deployment patterns to improve reliability, observability, and developer velocity on the ML platform. Requirements * 3+ years of experience ...
Machine Learning Engineer (AWS)
Tulsa, OK · On-site +1
Continuously explore new AWS services, ML frameworks, and deployment patterns to improve reliability, observability, and developer velocity on the ML platform. Requirements * 3+ years of experience ...
Machine Learning Engineer (AWS)
Tulsa, OK · On-site +1
Continuously explore new AWS services, ML frameworks, and deployment patterns to improve reliability, observability, and developer velocity on the ML platform. Requirements * 3+ years of experience ...
IT Director, Cloud Enablement & Transformation
Tulsa, OK · On-site
$180 - $250/hr
Experience with observability platforms, monitoring solutions, operational analytics, Site Reliability Engineering principles, incident management, problem management, change management, and ITIL ...
IT Director, Cloud Enablement & Transformation
Tulsa, OK · On-site
$180 - $250/hr
Experience with observability platforms, monitoring solutions, operational analytics, Site Reliability Engineering principles, incident management, problem management, change management, and ITIL ...
... observability platforms, monitoring solutions, operational analytics, Site Reliability Engineering principles, incident management, problem management, change management, and ITIL practices. • ...
... observability platforms, monitoring solutions, operational analytics, Site Reliability Engineering principles, incident management, problem management, change management, and ITIL practices. • ...
Observability Site Reliability Engineer information
What is the difference between Observability Site Reliability Engineer vs Monitoring Engineer?
| Aspect | Observability Site Reliability Engineer | Monitoring Engineer |
|---|---|---|
| Focus | Ensuring system reliability through observability, automation, and incident response | Implementing and managing monitoring tools and dashboards |
| Skills | Cloud platforms, scripting, incident management, observability tools | Monitoring tools, alerting systems, data analysis |
| Work Environment | DevOps teams, cloud infrastructure, large-scale systems | Operations teams, infrastructure monitoring |
While both roles involve system health, the Observability Site Reliability Engineer focuses on comprehensive system reliability using observability practices, whereas Monitoring Engineers primarily manage monitoring tools and alerts. The SRE role emphasizes automation, incident response, and system resilience, making it broader in scope.
What are popular job titles related to Observability Site Reliability Engineer jobs in Oklahoma?
For Observability Site Reliability Engineer jobs in Oklahoma, the most frequently searched job titles are:
What job categories do people searching Observability Site Reliability Engineer jobs in Oklahoma look for?
The top searched job categories for Observability Site Reliability Engineer jobs in Oklahoma are:
What cities in Oklahoma are hiring for Observability Site Reliability Engineer jobs?
Cities in Oklahoma with the most Observability Site Reliability Engineer job openings:
Manager of Cloud Platform Operations
Oklahoma City, OK • Hybrid
Full-time
Medical, Dental, Vision, Retirement, PTO
Re-posted yesterday
Love's Travel Stops rating
5.9
Based on 799 frontline employees who took The Breakroom Quiz
402nd of 734 rated retailers
Job description
Benefits: * Fuel Your Growth with Love's - company funded tuition assistance * Paid Time Off * 401(k) – 100% Match up to 5% * Medical/Dental/Vision Insurance after 30 days * Competitive Pay * Career Development
Role Summary
The Manager of Cloud Platform Operations leads a team of Cloud Platform Operations Engineers responsible for the stability, performance, and continuous improvement of the enterprise’s Azure- and AWS-based cloud platform and distributed edge compute environments. This hands-on leadership role ensures high system reliability, rapid incident response, and operational excellence in a fast‑paced, ticket‑driven environment.
The manager provides technical direction, operational oversight, and people leadership while driving automation, AI‑enabled operations, and scalable support practices. This role partners closely with other infrastructure, engineering, and security teams to ensure the cloud platforms are secure, cost‑optimized, and aligned with enterprise goals.
ESSENTIAL DUTIES
Team Leadership & Operational Management- Lead, mentor, and develop a team of Cloud Platform Operations Engineers, fostering a culture of ownership, accountability, and continuous improvement.
- Oversee daily operational activities, ensuring timely response to incidents, service requests, and operational tasks.
- Prioritize and assign work across a ticket‑driven environment with defined SLAs.
- Conduct performance reviews, coach team members, and support career development pathways.
- Ensure reliable operation of AWS and Azure infrastructure and edge compute environments.
- Provide escalation support for complex issues spanning compute, storage, networking, access, and hybrid integrations.
- Validate and approve changes, deployments, and infrastructure modifications to maintain platform stability.
- Oversee major incident response, ensuring rapid restoration of service and clear communication to stakeholders.
- Lead root cause analysis efforts and ensure corrective actions are implemented and documented.
- Identify systemic issues and drive long‑term remediation through automation, design improvements, or process changes.
- Champion an automation‑first culture, ensuring repeatable tasks are codified using Terraform, CloudFormation, or scripting.
- Guide the team in leveraging AI tools (LLMs, copilots, log analysis platforms) to accelerate diagnostics and operational workflows.
- Partner with engineering teams to develop intelligent runbooks, automated workflows, and AI‑assisted operational capabilities.
- Monitor platform health, performance, and cost utilization; drive proactive improvements.
- Implement cost‑control strategies and ensure efficient use of cloud resources.
- Support reliability engineering practices to improve uptime, resilience, and service continuity.
- Oversee support for edge compute and virtualization environments, ensuring connectivity and synchronization with cloud platforms.
- Guide troubleshooting efforts for hybrid infrastructure issues spanning on‑premise, edge, and cloud systems.
- Ensure the team maintains accurate runbooks, SOPs, and knowledge articles.
- Drive operational discipline, consistency, and adherence to best practices.
- Promote knowledge sharing and cross‑team collaboration across infrastructure, engineering, and security groups.
SKILLS & QUALIFICATIONS
Experience
- 6–10+ years of experience in cloud or infrastructure operations, including AWS and/or production environments.
- 2–4+ years of leadership experience (team lead, supervisor, or manager) in a technical operations setting.
- Demonstrated success leading teams in ticket‑driven environments with SLAs and incident management processes.
- Hands‑on experience with Infrastructure‑as‑Code (Terraform preferred) and automation tooling.
- Experience supporting hybrid or distributed environments, including edge compute or virtualization platforms.
- Familiarity with AI‑enabled operational tooling and practices.
Technical Skills
- AWS core services: EC2, S3, IAM, VPC, etc.
- Azure core services: Azure VMs, Blob storage, Azure AD, VNets, etc.
- Infrastructure as Code: Terraform, CloudFormation.
- Systems administration (Linux, Windows).
- Monitoring, observability, and incident response tooling.
- Networking fundamentals: DNS, routing, firewalls, connectivity troubleshooting.
- Automation/scripting (Python, Bash, or similar).
- Cloud security fundamentals: IAM, access controls, and best practices.
Soft Skills
- Strong people leadership with the ability to coach, mentor, and develop technical talent.
- Excellent decision‑making under pressure, especially during high‑impact incidents.
- Ability to translate complex technical issues into clear, actionable communication.
- Strategic mindset with a focus on operational excellence and continuous improvement.
- Highly collaborative, working effectively across infrastructure, engineering, and security teams.
- Strong organizational skills with the ability to manage competing priorities.
Education:
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field required.
- Master’s degree or relevant certifications (AWS, Microsoft, ITIL) preferred.
WORK ENVIRONMENT
- Prolonged sitting, some bending and stooping
- Eye strain (screen use)
- Manual dexterity sufficient to operate a computer keyboard and calculator
- Occasional lifting of up to 25 pounds
- Requires normal range of hearing and vision
- Additional hours may be necessary.
Note: The items identified above are representative of those commonly associated with this position but are not exhaustive. Employees may encounter additional or unforeseen responsibilities in the course of their duties.
- This job description should not be construed to imply that these requirements are the exclusive standards of the position. All employees may be required to follow any other instructions, cross train in other positions, and perform other duties as required by workloads.
Our Culture:
Fueling customers' journeys since 1964, innovation leads the way for this family-owned and operated business headquartered in Oklahoma City. With nearly 40,000 team members, travel stops are the core business along with products and services that provide value for professional drivers, fleets, traveling public, RVers, alternative energy and wholesale fuel customers. Giving back to communities and an inclusive workplace are hallmarks of the award-winning culture.
Love's is an Equal Opportunity Employer. Veterans encouraged to apply.
What Love's Travel Stops employees say
Pay
Benefits
Hours and flexibility
Workplace
Get the full story on Breakroom
About Love's Travel Stops
Sourced by ZipRecruiter
Industry
Gasoline stations
Company size
10,000+ Employees
Headquarters location
Oklahoma City, OK, US