1

Platform Monitor Jobs in Ridgewood, NY (NOW HIRING)

Platform Engineer

New York, NY ยท On-site

$150K - $200K/yr

The platform processes millions of clinical documents monthly across multi-tenant deployments in ... Take ownership of at least one customer deployment end-to-end - monitoring, alerting, incident ...

Own API versioning, rate limiting, SLA monitoring, and client onboarding * Security & Access ... Ensure the platform scales from a small number of pilot clients to 2,000+ users without ...

The team is also responsible for the production environment which requires knowledge with monitoring, alerting and on-call practices. The Platform team performs many functions whether that is ...

As a QIS Platform Analyst, you will operate at the intersection of Python engineering, data, AI and front office processes, with direct ownership of platform analytics, intelligent monitoring and ...

Sr. Platform Engineer

Manhattan, NY ยท On-site

$147K - $170K/yr

... monitoring, logging, tracing, alerting, and SLO practices to keep the developer platform stable and measurable. - Drive platform reliability and operational excellence through incident response, root ...

Sr. Platform Engineer

Manhattan, NY ยท On-site

$147K - $170K/yr

... monitoring, logging, tracing, alerting, and SLO practices to keep the developer platform stable and measurable. - Drive platform reliability and operational excellence through incident response, root ...

Sr. Platform Engineer

Manhattan, NY ยท On-site

$147K - $170K/yr

... monitoring, logging, tracing, alerting, and SLO practices to keep the developer platform stable and measurable. - Drive platform reliability and operational excellence through incident response, root ...

Monitor platform support Slack channels and address customer questions and issues* Stay up-to-date with the latest Google Workspace features and best practices, implementing updates and improvements ...

Dario is seeking a hands-on Platform Engineer with at least three years of experience to join our ... Manage, monitor, and optimize AWS cloud infrastructure. * Build and maintain CI/CD pipelines and ...

Implement data quality monitoring with traceability back to source systems * Define schemas ... Collaborate on platform architecture decisions and help establish engineering best practices

Implement data quality monitoring with traceability back to source systems * Define schemas ... Collaborate on platform architecture decisions and help establish engineering best practices

Platform Engineer

New York, NY ยท On-site +1

$150K - $250K/yr

... monitoring standards that translate into actionable insights for product teams downstream. You've ... founding Platform Engineer at a fintech company where your work directly determines what we can ...

We're hiring a Platform Engineer to help world-class financial institutions automate away their ... Ensure systems are resilient through monitoring, alerting, and automated recovery About You * 5+ ...

We're hiring a Platform Engineer to help world-class financial institutions automate away their ... Ensure systems are resilient through monitoring, alerting, and automated recovery About You * 5+ ...

Automate the deployment, scaling, and monitoring of Kubernetes clusters * Assist with developing ... Keep pace with emerging tools, techniques, and cloud platforms * Experience building solutions in ...

Showing results 21-40

Platform Monitor information

See Ridgewood, NY salary details

$33

$65

$97

How much do platform monitor jobs pay per hour?

As of Sep 15, 2026, the average hourly pay for platform monitor in Ridgewood, NY is $65.50, according to ZipRecruiter salary data. Most workers in this role earn between $51.68 and $75.58 per hour, depending on experience, location, and employer.

GCP Agentic Platform Support Lead

New York, NY โ€ข On-site

Contractor

Re-posted 3 days ago


Job description

Role : GCP Agentic Platform Support Lead

Location : New York, NY 10019 (Need local candidates/Hybrid)
Client: Persistent


Detailed JD:

The platform support lead will set the foundation and requirements for support on the GCP Data & AI platform. They will define standards for platform health, managing incident resolution, and executing routine maintenance to support the platform. They will develop GCP cloud logging and monitoring reports to support visibility across the platform.

Activities are comprised of:

1.       SLA & Reliability Reporting

1.       Establish the initial framework for tracking Mean Time to Repair (MTTR) and Mean Time Between Failures (MTBF)

2.       Configure self-service billing and uptime dashboards for Con Edison stakeholders

2.       Foundation, Maintenance & Optimization

1.       Develop and deploy the initial suite of Cloud Logging and Monitoring reports to establish platform visibility

2.       Monitor GCP billing for anomalies (e.g., BigQuery slot spikes) and implement tactical fixes to ensure budget adherence

3.       Build and maintain the "Golden Path" runbooks to ensure operational procedures are documented as they are established

3.       Platform Monitoring & Incident Management

1.       Conduct solo reviews of overnight batch processing logs (e.g., Cloud Composer/Dataflow) to verify completion and identify failures before business hours progress

2.       Receive and prioritize platform-related tickets; determine if issues stem from infrastructure, pipelines, or upstream sources

3.       Execute root cause analysis (RCA) and apply fixes for code-based failures, IAM errors, or configuration drifts

4.       Act as the primary technical point of contact for Google Cloud Support or Con Edison Source System teams (SAP, GIS) when issues are external to the platform

4.       Minor Enhancements (Capacity-Based

1.       Maintain a prioritized backlog of minor requests to be addressed only after platform stability and incidents are managed

2.       Within available bandwidth, execute minor schema updates, ingestion schedule tweaks, or IAM modifications

Workstream Deliverables:

1.       Operations Runbook: The definitive MS Word resource reflecting current operational procedures and recovery steps (MS Word)

2.       Integrated Health & Cost Reporting: Automated tracking of service uptime and GCP spend via Cloud Monitoring (Cloud Monitoring Reports)

3.       Unified Incident & RCA Logs: A centralized record of Critical/High severity incidents and their resolutions, stored in the agreed management tool (ServiceNow/Jira or similar)

4.       Recovery & Maintenance Code: Validated code merged into the repository for bug fixes and configuration updates, including detailed release notes (GCP Code)