1

Ai Infrastructure Jobs in San Ramon, CA (NOW HIRING)

We need an AI Infrastructure Lead to own the design and operation of our GPU cluster management layer, model serving pipeline, and low-latency routing system. You will work directly with the CTO and ...

AI Infrastructure Engineer

Fremont, CA · On-site

$126K - $165K/yr

* Design, build, and operate on-prem infrastructure that behaves like a cloud environment for internal teams, including AI/ML workloads * Own datacenter and infrastructure operations: compute, storage ...

AI Infrastructure Engineer

Fremont, CA · On-site

$126K - $165K/yr

* Design, build, and operate on-prem infrastructure that behaves like a cloud environment for internal teams, including AI/ML workloads * Own datacenter and infrastructure operations: compute, storage ...

AI Infrastructure Engineer

Fremont, CA · On-site

$117K - $154K/yr

* Design, build, and operate on-prem infrastructure that behaves like a cloud environment for internal teams, including AI/ML workloads * Own datacenter and infrastructure operations: compute, storage ...

Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct ...

AI Infrastructure Engineer

Sunnyvale, CA · On-site

$126K - $165K/yr

The function owns the reliability, scalability, and operability of Meshy\'s AI model serving stack, along with core engineering infrastructure. The team operates a conventional production ...

AI Infrastructure Engineer

San Francisco, CA · On-site

$126K - $166K/yr

Spellbrush, the world's leading generative AI studio behind niji・journey , is looking for an AI Infrastructure Engineer to join us in building out end-to-end ML infrastructure to run our models on ...

AI Infrastructure Engineer

San Francisco, CA · On-site

$126K - $166K/yr

Spellbrush, the world's leading generative AI studio behind niji・journey , is looking for an AI Infrastructure Engineer to join us in building out end-to-end ML infrastructure to run our models on ...

next page

Showing results 1-20

Ai Infrastructure information

See San Ramon, CA salary details

$31

$66

$97

How much do ai infrastructure jobs pay per hour?

As of Aug 11, 2026, the average hourly pay for ai infrastructure in San Ramon, CA is $66.14, according to ZipRecruiter salary data. Most workers in this role earn between $53.75 and $77.12 per hour, depending on experience, location, and employer.

What is the difference between Ai Infrastructure vs Data Engineer?

AspectAi InfrastructureData Engineer
Required CredentialsBachelor's in CS, Engineering, or related; knowledge of cloud platforms and AI toolsBachelor's in CS, Data Science, or related; programming and database skills
Work EnvironmentCloud environments, AI model deployment, infrastructure setupData pipelines, database management, data processing
Employer & Industry UsageTech companies, AI startups, cloud providersTech firms, finance, healthcare, e-commerce

Ai Infrastructure professionals focus on building and maintaining the hardware and software systems that support AI models, while Data Engineers develop and manage data pipelines and databases. Both roles require technical skills and often collaborate but serve different core functions within AI and data ecosystems.

How much do AI infrastructure engineers make?

AI infrastructure engineers typically earn between $100,000 and $150,000 annually, depending on experience, location, and company size. Senior roles or those with specialized skills in cloud platforms and hardware may earn higher salaries, often exceeding $180,000.

What are AI infrastructure jobs?

AI infrastructure jobs involve designing, building, and maintaining the hardware, software, and network systems necessary to support artificial intelligence applications. These roles often require knowledge of cloud computing, data centers, machine learning frameworks, and system optimization to ensure reliable and efficient AI model deployment and operation.

What are the key skills and qualifications needed to thrive in AI infrastructure?

To thrive in AI Infrastructure, you need expertise in software engineering, distributed systems, cloud platforms, and a solid understanding of machine learning workflows, often supported by degrees in computer science or related fields. Familiarity with tools like Kubernetes, Docker, Terraform, and cloud services (AWS, GCP, Azure), as well as experience with CI/CD pipelines and monitoring systems, is essential. Strong problem-solving abilities, effective communication, and adaptability help professionals excel in cross-functional teams and rapidly evolving environments. These skills and qualities are crucial for building scalable, reliable systems that power AI applications and support organizational innovation.

What are common challenges faced by professionals working in AI infrastructure roles, and how can they be addressed?

Professionals in AI Infrastructure roles often encounter challenges related to scalability, system reliability, and integration with existing IT environments. Managing rapidly growing datasets and ensuring seamless deployment of machine learning models can be complex, requiring robust automation and monitoring tools. Collaboration with data scientists, software engineers, and DevOps teams is critical to ensure infrastructure meets the evolving needs of AI projects. Staying updated with the latest cloud technologies and best practices can help address these challenges and drive successful AI implementations.

What is AI infrastructure?

AI infrastructure refers to the combination of hardware, software, and cloud-based solutions that support the development, deployment, and scaling of artificial intelligence applications. It includes components such as GPUs, CPUs, storage systems, networking, data management tools, and machine learning frameworks. The goal of AI infrastructure is to provide the computational power and resources needed to train, test, and run AI models efficiently, whether on-premises or in the cloud. Organizations invest in robust AI infrastructure to accelerate innovation, manage large datasets, and ensure the reliability of their AI systems.
What are popular job titles related to Ai Infrastructure jobs in San Ramon, CA? For Ai Infrastructure jobs in San Ramon, CA, the most frequently searched job titles are:
What job categories do people searching Ai Infrastructure jobs in San Ramon, CA look for? The top searched job categories for Ai Infrastructure jobs in San Ramon, CA are:
What cities near San Ramon, CA are hiring for Ai Infrastructure jobs? Cities near San Ramon, CA with the most Ai Infrastructure job openings:
Infographic showing various Ai Infrastructure job openings in San Ramon, CA as of August 2026, with employment types broken down into 73% Full Time, 23% Part Time, and 4% Contract. Highlights an 65% Physical, 3% Hybrid, and 32% Remote job distribution, with an average salary of $137,570 per year, or $66.1 per hour.

AI Infrastructure Systems Engineer

Together AI

San Francisco, CA • On-site

$190K - $270K/yr

Full-time

Medical

Re-posted 26 days ago


Job description

Build the infrastructure powering the next generation of AI.

At Together AI, you'll build and operate one of the world's largest GPU fleets used for frontier model training and inference. This isn't a traditional infrastructure role-we're looking for engineers who love building systems, automating everything, and solving problems at massive scale.

You'll Thrive Here If You:

  • Love building systems that replace repetitive operational work.
  • Think of infrastructure as a software engineering problem.
  • Enjoy solving hard problems with no existing playbook.
  • Care deeply about performance, reliability, and scale.
  • Want to build technology that powers frontier AI models.

Our mission is simple: build AI infrastructure that largely runs itself-where intelligent systems deploy, monitor, diagnose, optimize, and heal GPU fleets at massive scale. Every system you build will directly improve the speed, efficiency, and reliability of one of the world's most advanced AI compute platforms.

If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we'd love to talk.

Responsibilities
  • Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention.
  • Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation.
  • Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers.
  • Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators.
  • Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads.
  • Build internal platforms and developer tools that allow infrastructure to be managed through software-not manual operations.
  • Continuously improve deployment velocity, reliability, and operational efficiency through automation.
  • Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure.
Requirements
  • 3+ years building distributed systems, infrastructure platforms, or large-scale backend software.
  • Strong software engineering skills in Python, Go, or Rust.
  • Experience building platforms, automation systems, or developer infrastructure.
  • Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies.
  • Strong systems thinking with the ability to understand problems across hardware and software.
  • A passion for solving complex infrastructure challenges through software.
  • An automation-first mindset-if a task is repeated, your instinct is to build a system to eliminate it.
Bonus Experience
  • GPU infrastructure, CUDA, NCCL, NVLink/NVSwitch
  • InfiniBand or RoCE networking
  • Bare-metal provisioning and lifecycle management
  • Large-scale AI training or inference clusters
  • Hardware health monitoring and predictive failure detection
  • Distributed storage systems
  • AI agents and autonomous infrastructure operations
About Together AI

Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers and engineers in our journey in building the next generation AI infrastructure.

Compensation

We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $190,000 - $270,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.