1

Ai Reliability Engineer Jobs in Secaucus, NJ (NOW HIRING)

GCP Site Reliability Engineer Interview Mode: candidates local to Parsippany, Nj who can attend an ... Familiarity with Google BI and AI/ML tools (Looker, BigQuery ML, Vertex AI, etc.) Experience with ...

GCP Site Reliability Engineer

Parsippany, NJ · On-site

$57.25 - $76.25/hr

Role- GCP Site Reliability Engineer Duration 12 month Onsite at Parsippany, New Jersey 07054 in ... Familiarity with Google BI and AI/ML tools (Looker, BigQuery ML, Vertex AI, etc.) Experience with ...

GCP Site Reliability Engineer

Parsippany, NJ · On-site

$57.25 - $76.25/hr

Role- GCP Site Reliability Engineer Duration 12 month Onsite at Parsippany, New Jersey 07054 in ... Familiarity with Google BI and AI/ML tools (Looker, BigQuery ML, Vertex AI, etc.) Experience with ...

Site Reliability Engineer III

Jersey City, NJ · On-site

$62.25 - $82.75/hr

As a Site Reliability Engineer III at JPMorgan Chase within the Chief Data & Analytics Office (CDAO ... Uses enterprise-authorized AI capabilities within the work environment to accelerate incident ...

Site Reliability Engineer

New York, NY · On-site +1

$165K - $330K/yr

By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable ... THE ROLE As a Site Reliability Engineer at Baseten, you'll define and codify the gold standards of ...

Mid-Level Site Reliability Engineer

New York, NY · On-site

$62.25 - $82.75/hr

About InstaLILY InstaLILY is an AI products and infrastructure company that puts execution at the ... As a Mid-Level Site Reliability Engineer, you will help build and own meaningful pieces of our IDP ...

SRE

Manhattan, NY · On-site

$110 - $140/hr

By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable ... THE ROLE As a Site Reliability Engineer at Baseten, you'll define and codify the gold standards of ...

Lead Site Reliability Engineer

New York, NY

$62.25 - $82.75/hr

As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank ... Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e ...

Lead Site Reliability Engineer

Manhattan, NY · On-site

$62.75 - $83.25/hr

As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank ... Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e ...

AI is a core part of how you work - you've gone well beyond basic copilots and actively use AI ... Partner closely with engineering and product teams to maintain reliability without slowing ...

Site Reliability Engineer

New York, NY · On-site

$260K - $300K/yr

We are an applied AI lab building end-to-end software agents. We're the makers of Devin, the first AI software engineer. Our team is extremely talent-dense. Among our founding team, we have world ...

Showing results 21-40

Ai Reliability Engineer information

See Secaucus, NJ salary details

$62K

$119.9K

$143.4K

How much do ai reliability engineer jobs pay per year?

As of Sep 5, 2026, the average yearly pay for ai reliability engineer in Secaucus, NJ is $119,941.00, according to ZipRecruiter salary data. Most workers in this role earn between $104,200.00 and $131,200.00 per year, depending on experience, location, and employer.

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What are popular job titles related to Ai Reliability Engineer jobs in Secaucus, NJ?

For Ai Reliability Engineer jobs in Secaucus, NJ, the most frequently searched job titles are:

What job categories do people searching Ai Reliability Engineer jobs in Secaucus, NJ look for?

The top searched job categories for Ai Reliability Engineer jobs in Secaucus, NJ are:

What cities near Secaucus, NJ are hiring for Ai Reliability Engineer jobs?

Cities near Secaucus, NJ with the most Ai Reliability Engineer job openings:

Infographic showing various Ai Reliability Engineer job openings in Secaucus, NJ as of August 2026, with employment types broken down into 74% Full Time, 21% Part Time, 1% Temporary, and 4% Contract. Highlights an 65% Physical, 5% Hybrid, and 30% Remote job distribution, with an average salary of $119,941 per year, or $57.7 per hour.

Senior Site Reliability Engineer (SRE

Gov Services Hub

New York, NY • On-site, Remote

$62.25 - $82.75/hr

Full-time

Posted 9 days ago


Job description

Title: Senior Site Reliability Engineer (SRE)
Location: Remote

About
January

At
January, we’re transforming the lives of borrowers by bringing humanity to
consumer finance. Our data-driven products empower financial institutions to
streamline collections and help borrowers regain financial stability and
control over their lives. We’re not just expanding access to credit — we’re
restoring dignity and paving the way for millions to achieve financial freedom.

About
the Role

As a Senior
Site Reliability Engineer (SRE)
, you will establish SRE practices from the
ground up — ensuring reliability, scalability, and performance as January
scales from thousands to millions of borrowers. You’ll architect resilient
infrastructure, design modern observability solutions, and build sustainable
on-call processes that evolve with our rapid growth.

Your work
will directly address scaling challenges including database optimization, async
workflow infrastructure, and data pipeline reliability — enabling the
engineering team to ship confidently and efficiently.

Key
Responsibilities

  • Lead incident response and develop
    sustainable on-call practices, including runbooks, blameless postmortems,
    and continuous improvement to reduce MTTR.
  • Build and maintain self-service observability tools
    (Datadog, Prometheus, ELK) for proactive monitoring and troubleshooting.
  • Create and maintain Infrastructure as Code (IaC) using Terraform or CloudFormation for consistent, secure
    AWS environments.
  • Partner with development teams to architect
    resilient, scalable infrastructure
     for critical components like
    databases, networking, async workflows, and data pipelines.
  • Design and implement robust CI/CD pipelines (GitHub
    Actions) with advanced deployment strategies (blue/green, canary).
  • Drive best practices in reliability and
    performance early in the design phase to future-proof January’s systems.

Required
Skills & Experience

  • Proven experience leading incident response and
    postmortem processes for high-availability production systems.
  • Deep expertise in designing highly available
    architectures
     (EC2, Fargate, auto-scaling, health checks,
    graceful degradation).
  • Strong experience with AWS cloud infrastructure and IaC tools (Terraform, CloudFormation).
  • Hands-on experience with CI/CD automation using GitHub Actions or equivalent tools.
  • Proficiency in observability and monitoring stacks (Datadog,
    Prometheus, ELK
    ).
  • Solid scripting/programming skills in Python (for
    automation, tooling, and debugging).
  • Excellent communication and documentation skills, with
    the ability to collaborate across engineering and platform teams.




Requirements

Tools
& Technologies

  • Cloud: AWS
  • IaC: Terraform,
    CloudFormation
  • CI/CD: GitHub
    Actions
  • Monitoring: Datadog,
    Prometheus, ELK
  • Languages: Python
  • Infrastructure: EC2,
    Fargate

Additional
Details

  • Remote role (NYC-based preferred for hybrid
    collaboration).
  • Opportunity to build and own the entire SRE practice for
    a growing FinTech startup.
  • Fast-paced, innovative environment working on
    AI-forward consumer finance products.