1

Observability Site Reliability Engineer Jobs in California

SME SRE Observability

Fremont, CA · On-site

$62.50 - $83.25/hr

Info Way Solutions is seeking an experienced Subject Matter Expert (SME) in Site Reliability Engineering (SRE) with a strong focus on Observability. The ideal candidate will be responsible for ...

Site Reliability Engineer

San Francisco, CA · Remote

$67.25 - $89.25/hr

The SRE defines and enforces SLOs aligned with contractual SLAs, builds the observability stack, and leads major incidents to resolution. Key responsibilities * Define and operate Service Level ...

Build and maintain observability infrastructure - logging, metrics, tracing, and alerting - so the ... Help define and build the SRE team as the company scales - this is a foundational hire with a path ...

Site Reliability Engineer (SRE)

San Francisco, CA · On-site

$67.25 - $89.25/hr

As a Site Reliability Engineer (SRE), you will be responsible for ensuring the reliability ... observability tools, including monitoring, logging, and alerting (e.g., Prometheus, Grafana ...

Site Reliability Engineer

Santa Clara, CA · On-site

$230K - $250K/yr

Build and maintain observability infrastructure - logging, metrics, tracing, and alerting - so the ... Help define and build the SRE team as the company scales - this is a foundational hire with a path ...

As a Site Reliability Engineer, you will strengthen infrastructure, optimize tooling, deepen observability, streamline incident response, and elevate reliability standards. These actions empower ...

About the DevOps / SRE Team The DevOps/SRE team at Obsidian ensures that engineering excellence ... Address complex challenges around scalability, reliability, observability, and cost efficiency

Site Reliability Engineer

San Francisco, CA · On-site

$67.25 - $89.25/hr

... Site Reliability Engineer to own infrastructure and reliability. This role involves designing ... Implement and improve observability tools, logging, and monitoring solutions to identify and ...

Site Reliability Engineer

Mountain View, CA · Hybrid

$67.25 - $89.25/hr

Site Reliability Engineer Onsite- Bay Area, CA Skills Relevant Skills and Experience What You'll Do ... Monitor system health and performance using Grafana and other observability tools. * Ensure high ...

The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the ... Develop and maintain observability tools , including monitoring, logging, and alerting (e.g ...

Principal III, SRE

Torrance, CA · Hybrid

$59.75 - $79.50/hr

The SRE team consists of: SRE Engineers Deployment Automation Incident Response and Postmortem Analysis Observability and MonitoringThis role will drive the adoption of best practices in multi-cloud ...

Principal III, SRE

Torrance, CA · Hybrid

$59.75 - $79.50/hr

The SRE team consists of: SRE Engineers Deployment Automation Incident Response and Postmortem Analysis Observability and Monitoring This role will drive the adoption of best practices in multi-cloud ...

Principal III, SRE

Torrance, CA · On-site

$59.75 - $79.50/hr

The SRE team consists of: • SRE Engineers • Deployment Automation • Incident Response and Postmortem Analysis • Observability and Monitoring This role will drive the adoption of best ...

Site Reliability Engineer

Newport Beach, CA · On-site

$61.25 - $81.50/hr

They are seeking a Site Reliability Engineer to support and maintain the service quality of their ... around scalability, reliability, observability, and cost efficiency • Collaborate with ...

next page

Showing results 1-20

Observability Site Reliability Engineer information

What is the difference between Observability Site Reliability Engineer vs Monitoring Engineer?

AspectObservability Site Reliability EngineerMonitoring Engineer
FocusEnsuring system reliability through observability, automation, and incident responseImplementing and managing monitoring tools and dashboards
SkillsCloud platforms, scripting, incident management, observability toolsMonitoring tools, alerting systems, data analysis
Work EnvironmentDevOps teams, cloud infrastructure, large-scale systemsOperations teams, infrastructure monitoring

While both roles involve system health, the Observability Site Reliability Engineer focuses on comprehensive system reliability using observability practices, whereas Monitoring Engineers primarily manage monitoring tools and alerts. The SRE role emphasizes automation, incident response, and system resilience, making it broader in scope.

What are popular job titles related to Observability Site Reliability Engineer jobs in California? For Observability Site Reliability Engineer jobs in California, the most frequently searched job titles are:
What job categories do people searching Observability Site Reliability Engineer jobs in California look for? The top searched job categories for Observability Site Reliability Engineer jobs in California are:
What cities in California are hiring for Observability Site Reliability Engineer jobs? Cities in California with the most Observability Site Reliability Engineer job openings:
Infographic showing various Observability Site Reliability Engineer job openings in California as of July 2026, with employment types broken down into 70% Full Time, and 30% Contract. Highlights an 85% In-person, 10% Hybrid, and 5% Remote job distribution.

SME SRE Observability

Info Way Solutions

Fremont, CA • On-site

$62.50 - $83.25/hr

Full-time

Re-posted 8 days ago


Job description

Job Summary:
Info Way Solutions is seeking an experienced Subject Matter Expert (SME) in Site Reliability Engineering (SRE) with a strong focus on Observability. The ideal candidate will be responsible for designing, implementing, and optimizing observability frameworks to ensure high system reliability, performance, and scalability in a production environment.
Responsibilities:
• Lead the design and implementation of observability solutions including metrics, logging, and tracing.
• Act as an SME for SRE best practices, ensuring system reliability, availability, and performance.
• Develop and maintain dashboards, alerts, and monitoring strategies.
• Collaborate with development, DevOps, and infrastructure teams to improve system visibility.
• Perform root cause analysis (RCA) and drive incident resolution.
• Optimize system performance and reliability through proactive monitoring.
• Implement automation to improve operational efficiency and reduce manual intervention.
• Define and track SLIs, SLOs, and SLAs.
Qualifications:
Required:
• Strong experience in Site Reliability Engineering (SRE) concepts and practices.
• Deep expertise in Observability tools (e.g., Prometheus, Grafana, ELK Stack, Datadog, Splunk, or similar).
• Experience with cloud platforms (AWS, Azure, or GCP).
• Proficiency in scripting/programming (Python, Bash, or similar).
• Hands-on experience with monitoring, alerting, and logging frameworks.
• Strong troubleshooting and performance tuning skills.
• Experience with CI/CD pipelines and automation tools.
Preferred:
• Experience working in high-availability, distributed systems.
• Knowledge of containerization and orchestration tools (Docker, Kubernetes).
• Prior experience as an SRE SME or Lead.
Company:
Founded and incorporated in 2012 , Info Way Solutions is an IT services and consulting company headquartered in Fremont , CA. Founded in 2012, the company is headquartered in Fremont, USA, with a team of 501-1000 employees. The company is currently Late Stage.