1

Ai Reliability Engineer Jobs in Virginia (NOW HIRING)

As a Site Reliability Engineer, you will be focused on supporting a specific, highly available, and ... Are comfortable using AI-assisted engineering tools such as Codex, Claude Code, and Cursor to ...

Staff Site Reliability Engineer

Reston, VA

$59.25 - $78.75/hr

The Site Reliability Engineering team drives reliability strategy, elevates engineering standards ... Hands-on experience designing and integrating AI/ML-powered solutions into cloud-native platforms ...

Staff Site Reliability Engineer

Reston, VA · On-site

$59.25 - $78.75/hr

The Site Reliability Engineering team drives reliability strategy, elevates engineering standards ... Hands-on experience designing and integrating AI/ML-powered solutions into cloud-native platforms ...

Staff Site Reliability Engineer

Reston, VA · On-site

$59.25 - $78.75/hr

The Site Reliability Engineering team drives reliability strategy, elevates engineering standards ... Hands-on experience designing and integrating AI/ML-powered solutions into cloud-native platforms ...

Senior Site Reliability Engineer

Tysons, VA · On-site

$57.50 - $76.50/hr

AI at Cvent: Leading the Future Are you ready to shape the future of work at the intersection of ... As a Senior Site Reliability Engineer, you will use your advanced development and operations ...

As a Senior Site Reliability Engineer at Commence, you will own the reliability, scalability, and ... Exposure to AI/ML infrastructure and the reliability challenges unique to model serving.

Site Reliability Engineer - CTJ - Poly

Reston, VA · On-site

$59.25 - $78.75/hr

You will gain end-to-end ownership experience across the service lifecycle while helping shape the future of reliability engineering through automation, data-driven operations, and AI-enabled ...

Showing results 21-40

Ai Reliability Engineer information

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.

What are popular job titles related to Ai Reliability Engineer jobs in Virginia?

For Ai Reliability Engineer jobs in Virginia, the most frequently searched job titles are:

What job categories do people searching Ai Reliability Engineer jobs in Virginia look for?

The top searched job categories for Ai Reliability Engineer jobs in Virginia are:

What cities in Virginia are hiring for Ai Reliability Engineer jobs?

Cities in Virginia with the most Ai Reliability Engineer job openings:

Infographic showing various Ai Reliability Engineer job openings in Virginia as of August 2026, with employment types broken down into 72% Full Time, 24% Part Time, and 4% Contract. Highlights an 67% Physical, 3% Hybrid, and 30% Remote job distribution.

Site Reliability Engineer - Networking

Cisco

Herndon, VA

$58.50 - $78/hr

Full-time

Medical, Dental, Vision, Life, Retirement, PTO

Posted 26 days ago


Cisco Systems rating

8.0

Company rating: 8.0 out of 10

Based on 42 frontline employees who took The Breakroom Quiz

58th of 157 rated electronics manufacturers


Job description

The application window is expected to close on: 08/03/2026

Job posting may be removed earlier if the position is filled or if a sufficient number of applications are received.

The successful applicant may be performing work in FedRAMP High or IL-5 environments, and therefore, must be a U.S. Person (i.e. U.S. citizen, U.S. national, lawful permanent resident, asylee, or refugee). This position may also perform work that the U.S. government has specified can only be performed by a U.S. citizen on U.S. soil.


The successful applicant may be performing work in FedRAMP High or IL-5 environments, and therefore, must be a U.S. Person (i.e. U.S. citizen, U.S. national, lawful permanent resident, asylee, or refugee). This position may also perform work that the U.S. government has specified can only be performed by a U.S. citizen on U.S. soil.

Meet the Team

The Cisco Meraki cloud supports millions of customer devices from 8 physical data centers and several cloud regions around the world. Cisco Meraki's customer base has grown by a factor of 2-3 every year, serving billions of HTTP requests per day globally. Our customers depend on our products to run their critical infrastructure of network switches, security appliances, wireless APs and security cameras.
As SREs at Cisco Meraki, we are responsible for building and growing the cloud that supports these customers and their networks. As a Site Reliability Engineer, you will be focused on supporting a specific, highly available, and very secure production environment. You will analyze the reliability of our environments, use your networking and coding skills to improve the way we operate, and debug complex events as part of a 24/7 on-call rotation.
This role will work with BGP, AWS Cloud Networking components, the Cisco Nexus family of switches and Cisco Firepower appliances.

Your Impact

Example projects of a Site Reliability Engineer (Network):
* Crafting the network architecture for a hybrid cloud design.
* Improving agility by automating endpoint network connectivity.
* Creating the best connectivity solution by abstracting network functions from hardware.
* Developing comprehensive monitoring tools that provide transparency into the performance and reliability of our network infrastructure.
* Full-stack troubleshooting of production issues (application, system and network) through to root cause analysis and implementation of preventative measures.
* Design, implementation and management of an overlay network to support 1000's of containers.

Minimum Qualifications

* Bachelors + 7 years of related experience, or Masters + 4 years of related experience, or PhD + 1 year of related experience.

You are an ideal candidate if you:

* Can demonstrate experience in designing, deploying and operating mid to large scale enterprise or cloud environments.
* Feel comfortable with scripting or coding with languages like Python, Bash, Ruby, Go
* Know your way around *nix systems and have interests that span beyond routers and switches
* Daringly jump into other people's solutions and source code to seek a problem.
* Like to Automate all the things!
* You care and empathize with the customer experience and can support an externally facing production environment.
* Are passionate about building complex solutions working alongside other engineering teams
* Empathize with coworkers and have a positive influence on others.
* Are comfortable using AI-assisted engineering tools such as Codex, Claude Code, and Cursor to support coding, code reviews, troubleshooting, testing, documentation, and incident investigation.
* You have experience with: BGP, OSPF, IPv6, Network Security, DMVPN, MACSec, Grafana, Splunk, advanced Unix/Linux, system administration, scripting/Bash/Ruby/Python, project management experience, AWS/Azure, Docker, K8s, Ansible, REST APIs, TLS or Cloud/ISP/Telco exposure.

Preferred Qualifications

* Understand agent-based workflows, including reusable skills and MCP integrations, and can critically evaluate AI-generated output for accuracy, security, reliability, and maintainability.
* Keywords: Network Engineering, Production Engineering, SRE, Site Reliability Engineering, DevOps, CCNP, CCIE, JNCP, JNCIE

Why Cisco?

At Cisco, we're revolutionizing how data and infrastructure connect and protect organizations in the AI era - and beyond. We've been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint.

Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you'll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere.

We are Cisco, and our power starts with you.

Message to applicants applying to work in the U.S. and/or Canada:The starting salary range posted for this position is $167,700.00 to $245,200.00 and reflects the projected salary range for new hires in this position in U.S. and/or Canada locations, not including incentive compensation*, equity, or benefits.

Individual pay is determined by the candidate's hiring location, market conditions, job-related skillset, experience, qualifications, education, certifications, and/or training. The full salary range for certain locations is listed below. For locations not listed below, the recruiter can share more details about compensation for the role in your location during the hiring process.

U.S. employees are offered benefits, subject to Cisco's plan eligibility rules, which include medical, dental and vision insurance, a 401(k) plan with a Cisco matching contribution, paid parental leave, short and long-term disability coverage, and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible to receive grants of Cisco restricted stock units, which vest following continued employment with Cisco for defined periods of time.

U.S. employees are eligible for paid time away as described below, subject to Cisco's policies:

  • 10 paid holidays per full calendar year, plus 1 floating holiday for non-exempt employees

  • 1 paid day off for employee's birthday, paid year-end holiday shutdown, and 4 paid days off for personal wellness determined by Cisco

  • Non-exempt employees** receive 16 days of paid vacation time per full calendar year, accrued at rate of 4.92 hours per pay period for full-time employees

  • Exempt employees participate in Cisco's flexible vacation time off program, which has no defined limit on how much vacation time eligible employees may use (subject to availability and some business limitations)

  • 80 hours of sick time off provided on hire date and each January 1st thereafter, and up to 80 hours ofunused sick timecarried forwardfrom one calendar yearto the next

  • Additional paid time away may be requested to deal with critical or emergency issues for family members

  • Optional 10 paid days per full calendar year to volunteer

For non-sales roles, employees are also eligible to earn annual bonuses subject to Cisco's policies.

Employees on sales plans earn performance-based incentive pay on top of their base salary, which is split between quota and non-quota components, subject to the applicable Cisco plan. For quota-based incentive pay, Cisco typically pays as follows:

  • .75% of incentive target for each 1% of revenue attainment up to 50% of quota;

  • 1.5% of incentive target for each 1% of attainment between 50% and 75%;

  • 1% of incentive target for each 1% of attainment between 75% and 100%; and

  • Once performance exceeds 100% attainment, incentive rates are at or above 1% for each 1% of attainment with no cap on incentive compensation.

For non-quota-based sales performance elements such as strategic sales objectives, Cisco may pay 0% up to 125% of target. Cisco sales plans do not have a minimum threshold of performance for sales incentive compensation to be paid.

The applicable full salary ranges for this position, by specific state, are listed below:

New York City Metro Area:

$167,700.00 - $282,000.00

Non-Metro New York state & Washington state:

$149,100.00 - $250,900.00

* For quota-based sales roles on Cisco's sales plan, the ranges provided in this posting include base pay and sales target incentive compensation combined.

** Employees in Illinois, whether exempt or non-exempt, will participate in a unique time off program to meet local requirements.


What Cisco Systems employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Cisco Systems logo

About Cisco Systems

Sourced by ZipRecruiter

Cisco Systems, a global tech titan based in San Jose, CA, US, operates in the information technology and services industry. Founded in 1984, the company was derived from a project between two computer scientists from Stanford University. They aimed to connect different networks of computer systems at the university, resulting in the first multi-protocol router, and subsequently, the birth of Cisco. As an industry-leading manufacturer of networking hardware and telecommunications equipment, Cisco's product and services range includes routers, switches, firewall devices, and telecommunication technology. The company's mission, "to shape the future of the Internet by creating unprecedented value and opportunity for our customers, employees, investors, and ecosystem partners," is a testament to its pursuit of technology-forward innovation and customer satisfaction.

Industry

Computer and computer peripheral equipment and software wholesalers

Company size

10,000+ Employees

Headquarters location

San Jose, CA, US

Year founded

1984

Social media