1

Ai Reliability Engineer Jobs in Texas (NOW HIRING)

Site Reliability Engineer

Austin, TX · On-site

$130 - $180/hr

Austin Reporting To: SRE Manager Description Are you an SRE ready to grow your impact at the world's largest AI cloud-native physical security company? Following the merger of Brivo and Eagle Eye ...

Quality & Reliability Engineer

Austin, TX · On-site

$101K - $127K/yr

Exowatt is revolutionizing the energy landscape for the AI era with our groundbreaking P3 system ... We're hiring a Quality & Reliability Engineer to own both the quality of our supply chain and ...

Quality & Reliability Engineer

Austin, TX

$101K - $127K/yr

Exowatt is revolutionizing the energy landscape for the AI era with our groundbreaking P3 system ... We're hiring a Quality & Reliability Engineer to own both the quality of our supply chain and ...

Quality & Reliability Engineer

Austin, TX

$101K - $127K/yr

Exowatt is revolutionizing the energy landscape for the AI era with our groundbreaking P3 system ... We're hiring a Quality & Reliability Engineer to own both the quality of our supply chain and ...

Site Reliability Engineer

Westlake, TX · On-site +1

$54.75 - $72.75/hr

We are currently seeking a Site Reliability Engineer to join our team in Westlake, Texas (US-TX), ... We are one of the world's leading AI and digital infrastructure providers, with unmatched ...

Site Reliability Engineer

Westlake, TX · On-site +1

$54.75 - $72.75/hr

We are currently seeking a Site Reliability Engineer to join our team in Westlake, Texas (US-TX), ... We are one of the world's leading AI and digital infrastructure providers, with unmatched ...

As a Site Reliability Engineer III at JPMorgan Chase within the Chief Data & Analytics Office (CDAO) AI/ML & Data Platforms team, you will solve complex and broad business problems with simple and ...

Site Reliability Engineer

Westlake, TX · On-site

$54.75 - $72.75/hr

We are currently seeking a Site Reliability Engineer to join our team in Westlake, Texas (US-TX), ... We are one of the world's leading AI and digital infrastructure providers, with unmatched ...

Staff Cyber Site Reliability Engineer (SRE)

Dallas, TX · On-site

$56.50 - $75/hr

Experience with AI/ML and working knowledge of LLMs is a meaningful differentiator. Position Responsibilities As a Staff Cyber SRE, you will: * Own Reliability Engineering: Define and drive ...

Senior Site Reliability Engineer

Austin, TX · On-site

$56.50 - $75/hr

But we're an AI-native company, and we expect you to use AI-assisted development (Claude Code) as a ... Partner closely with engineering on reliability reviews and architecture decisions Requirements * 5 ...

Showing results 21-40

Ai Reliability Engineer information

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.
What cities in Texas are hiring for Ai Reliability Engineer jobs? Cities in Texas with the most Ai Reliability Engineer job openings:
Infographic showing various Ai Reliability Engineer job openings in Texas as of August 2026, with employment types broken down into 78% Full Time, and 22% Contract. Highlights an 86% In-person, and 14% Remote job distribution.

Senior AI Ops & Incident/Site Reliability Engineer

Perficient, Inc.

Plano, TX • Hybrid

$53.25 - $70.75/hr

Full-time

Posted 18 days ago


Job description

We currently have a career opportunity for a NOC AI-Ops Engineer to join our team located in Austin, TX. This is a hybrid role, 3 days a week in office.

Job Overview:

We are seeking a Senior AIOps and Incident/Site Reliability Engineer to lead incident management, operational resilience, and intelligent automation initiatives across enterprise technology environments. This role will partner with Network Operations Center (NOC), Infrastructure Operations, Cloud Engineering, DevOps, and Application Support teams to proactively detect, respond to, and prevent technology incidents. The ideal candidate combines hands-on incident management expertise with experience implementing observability, automation, and AI-driven operational solutions to improve system reliability, reduce operational overhead, and enhance customer experience. The candidate should possess deep expertise in AIOps, ITSM, ITIL, SRE, Incident Management, Cloud Operations, and Enterprise Infrastructure.

This is a hybrid role requiring on-site attendance up to three days per week. Candidates must reside within approximately one hour commuting distance of one of the following office locations: Fort Mill, SC, Austin, TX, Boston, MA, New York, NY, Tempe, AZ, or San Diego, CA.

Perficient is always looking for the best and brightest talent and we need you! We're a quickly-growing, global digital consulting leader, and we're transforming the world's largest enterprises and biggest brands. You'll work with the latest technologies, expand your skills, and become a part of our global community of talented, diverse, and knowledgeable colleagues.

Perficient is the global AI and technology consulting firm disrupting the traditional consulting model. Powered by our 7,000+ advisors, engineers, and designers, Perficient implements AI-first solutions that break conventions and deliver outcomes that matter. Proudly serving clients that represent the world's most innovative brands, and in collaboration with our powerful technology partner ecosystem, we bring deep industry expertise and data-driven design to redefine how businesses run and succeed. Perficient is different. For real. Learn more at perficient.com.
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field (or equivalent experience).
  • 8+ years of experience in IT Operations, Site Reliability Engineering, Infrastructure Operations, Network Operations, or Production Support environments.
  • 5+ years of experience leading incident management, operational transformation, or reliability engineering initiatives.
  • Strong experience with:  
    • Site Reliability Engineering (SRE)
    • IT Service Management (ITSM)
    • ITIL Framework
    • Incident, Problem, Change, and Event Management
    • Network Operations Center (NOC)
    • Infrastructure Operations
    • Service Desk Operations
    • Application Production Support
    • Cloud Platforms (AWS, Azure, or GCP)
    • DevOps Practices and Toolchains
  • Hands-on experience with Dynatrace, monitoring platforms, and observability solutions.
  • Experience using ServiceNow for ticketing, workflow automation, and service management.
  • Strong understanding of infrastructure, networking, cloud architecture, and enterprise application ecosystems.
  • Proven experience conducting root cause analysis and implementing preventive controls.
  • Experience leading enterprise AIOps implementations.
  • Experience building AI-powered operational agents and intelligent automation solutions.
  • Certifications such as:  
    • ITIL Foundation or ITIL Managing Professional
    • Certified Site Reliability Engineer (SRE)
    • AWS, Azure, or Google Cloud certifications
    • ServiceNow certifications
  • Experience with workflow orchestration and enterprise automation platforms.
  • Familiarity with predictive analytics, machine learning operations, and autonomous operations frameworks. 
     

ADDITIONAL INFORMATION

Perficient, Inc. proudly provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, gender, sexual orientation, national origin, age, disability, genetic information, marital status, amnesty, or status as a protected veteran in accordance with applicable federal, state and local laws. Perficient, Inc. complies with applicable state and local laws governing nondiscrimination in employment in every location in which the company has facilities. This policy applies to all terms and conditions of employment, including, but not limited to, hiring, placement, promotion, termination, layoff, recall, transfer, leaves of absence, compensation, and training. Perficient, Inc. expressly prohibits any form of unlawful employee harassment based on race, color, religion, gender, sexual orientation, national origin, age, genetic information, disability, or covered veterans. Improper interference with the ability of Perficient, Inc. employees to perform their expected job duties is absolutely not tolerated.

Disability Accommodations: Perficient is committed to providing a barrier-free employment process with reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If you need assistance or accommodation due to a disability, please contact us.

Applications will be accepted until the position is filled or the posting is removed.

The salary range for this position takes into consideration a variety of factors, including but not limited to skill sets, level of experience, applicable office location, training, licensure and certifications, and other business and organizational needs. The new hire salary range displays the minimum and maximum salary targets for this position across all US locations, and the range has not been adjusted for any specific state differentials. It is not typical for a candidate to be hired at or near the top of the range for their role, and compensation decisions are dependent on the unique facts and circumstances regarding each candidate. A reasonable estimate of the current salary range for this position is $93,600.00 to $170,640.00. Please note that the salary range posted reflects the base salary only and does not include benefits or any potential equity or variable bonus programs. Information regarding the benefits available for this position are in our benefits overview.

Disclaimer:  The above statements are not intended to be a complete statement of job content, rather to act as a guide to the essential functions performed by the employee assigned to this classification.  Management retains the discretion to add or change the duties of the position at any time. 

#LI-GS1 #LI-AIFirst

  • Incident & Recovery Management
  • Monitor, document, and analyze major incident response efforts and service recovery activities.
  • Serve as a senior escalation point for Tier 1 and Tier 2 operational incidents.
  • Conduct incident reviews, root cause analysis, and corrective action planning.
  • Improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR).
  • Site Reliability Engineering
  • Implement SRE practices to improve platform reliability, scalability, and resiliency.
  • Define and monitor SLAs, SLOs, and operational KPIs.
  • Develop proactive reliability and availability strategies.
  • AIOps & Automation
  • Implement AIOps solutions to automate incident detection, diagnosis, remediation, and prevention.
  • Build and optimize AI-powered operational agents and self-healing workflows.
  • Reduce operational effort through intelligent automation.
  • Observability & Monitoring
  • Lead enterprise monitoring initiatives using Dynatrace and related observability platforms.
  • Improve visibility across cloud, infrastructure, applications, and user experiences.
  • Enable predictive monitoring and anomaly detection.
  • ITSM & Service Operations
  • Develop and enhance incident, problem, change, and event management frameworks aligned with ITIL and ITSM best practices.  
  • Leverage ServiceNow workflow automation to improve service delivery.
  • Cross-Functional Leadership
  • Partner with Infrastructure, DevOps, Cloud, Security, Application Development, and NOC teams.
  • Mentor operational teams and promote an automation-first culture.