1

Big Data Infrastructure Engineer Jobs in California

Helix AI Engineer, Data Infrastructure

San Jose, CA · On-site

$126K - $165K/yr

They are seeking an experienced Data Infrastructure Engineer to enhance their AI data infrastructure by building tools and software components for managing robot data and cloud resources.

Helix AI Engineer, Data Infrastructure

San Jose, CA · On-site

$126K - $165K/yr

They are seeking an experienced Data Infrastructure Engineer to enhance their AI data infrastructure by building tools and software components for robot data management and supporting AI researchers ...

Big Data Engineer

Los Angeles, CA · On-site

$60 - $79.50/hr

Company Description Intelliswift Software, Inc As a Big Data Engineer, you will be an integral member of our threat intelligence service, i.e. auto focus, team responsible for architecture, design ...

Our Helix team is looking for an experienced Data Infrastructure Engineer, to take our AI data infrastructure to the next level. This role is focused on building tools and software components that ...

Our Helix team is looking for an experienced Data Infrastructure Engineer, to take our AI data infrastructure to the next level. This role is focused on building tools and software components that ...

Showing results 41-60

Big Data Infrastructure Engineer information

See California salary details

$45.9K

$125.4K

$179.6K

How much do big data infrastructure engineer jobs pay per year?

As of Sep 12, 2026, the average yearly pay for big data infrastructure engineer in California is $125,402.00, according to ZipRecruiter salary data. Most workers in this role earn between $106,100.00 and $139,200.00 per year, depending on experience, location, and employer.

What is a big data infrastructure engineer?

Big Data Infrastructure Engineers are IT professionals responsible for designing, building, and maintaining the systems and environments that process and store large volumes of data. They work with technologies such as Hadoop, Spark, and cloud platforms to ensure data pipelines are efficient, scalable, and reliable. These engineers also manage data storage, security, and high-performance computing resources, enabling organizations to analyze and leverage big data for business insights.

What are the key skills and qualifications needed to thrive as a big data infrastructure engineer?

To thrive as a Big Data Infrastructure Engineer, you need expertise in distributed systems, data architecture, and programming languages like Java, Scala, or Python, typically backed by a relevant degree in computer science or engineering. Proficiency with big data tools such as Hadoop, Spark, Kafka, and experience with cloud platforms (AWS, Azure, or GCP) and containerization technologies are crucial, as are certifications in these areas. Strong problem-solving, teamwork, and communication skills set top performers apart in this role. These skills enable efficient design, deployment, and maintenance of scalable data systems that support business intelligence and analytics needs.

What are some common challenges big data infrastructure engineers face when ensuring system scalability and reliability?

Big Data Infrastructure Engineers often encounter challenges related to optimizing system performance as data volumes grow rapidly. Ensuring that the infrastructure can scale horizontally without downtime requires careful planning and implementation of distributed computing frameworks. Additionally, maintaining data integrity and fault tolerance in complex, multi-node environments can be demanding, requiring constant monitoring and quick responses to system failures. Collaborating closely with data engineers and DevOps teams is essential to proactively address these challenges and implement robust solutions.

What is the difference between Big Data Infrastructure Engineer vs Data Engineer?

AspectBig Data Infrastructure EngineerData Engineer
Primary FocusDesigning, building, and maintaining big data infrastructure and pipelinesDeveloping and managing data pipelines, databases, and data models
Skills & CertificationsHadoop, Spark, cloud platforms, Linux, scriptingSQL, Python, ETL tools, cloud services
Work EnvironmentData centers, cloud environments, big data platformsData warehouses, cloud platforms, analytics teams
Industry UsageTech, finance, healthcare, any data-heavy industryTech, finance, retail, analytics-focused companies

While both roles involve working with data, Big Data Infrastructure Engineers focus on building and maintaining the infrastructure for processing large datasets, whereas Data Engineers develop data pipelines and models to enable data analysis. Understanding these distinctions helps in choosing the right career path or job search focus.

Principal Data Engineer (Python, AWS, Redshift, EMR, Airflow, Databricks, Big data) | Contract [...]

Irvine, CA • On-site

$122K - $147K/yr

Other

Posted 26 days ago


Key responsibilities

  • Lead the creation of data environments and data sets to serve various data users.

  • Own product data sets from definition through to production deployment.

  • Design, develop, and deploy scalable, reliable data pipelines and big data infrastructure.


Job description

  • Irvine, CA
Principal Data Engineer (Python, AWS, Redshift, EMR, Airflow, Databricks, Big data) | Contract to Hire | Irvine, CA (Hybrid) |

Title: Principal Data Engineer (Python, AWS, Redshift, EMR, Airflow, Databricks, Big-data)

Location: Irvine, CA (Hybrid)

Position Type: Contract to Hire

Responsibilities
  • As a Data Software Engineer Sr. Staff, you will be responsible for all aspects of data acquisition, data transformation, analytics scheduling and operationalization to drive high-visibility, cross-division outcomes. Investigate, evaluate, test and recommend technical solutions for future systems. They will support software developers, database architects, data scientists on data initiatives and will contribute optimal data delivery architecture.
What you will be doing
  • Data Operations
  • Lead the creation of data environments and/or data sets to serve a wide range of data users, including but not limited to Data Scientists, Data Analysts, Business Analysts etc.
  • Perform offline analysis of large data sets using components of a big data software ecosystem.
  • Validate the solution of root cause analysis escalated by various technical staff in multiple organizations and with differing levels of expertise.
  • Investigate, evaluate, test and recommend technical solutions for future systems.
Data Management
  • Own product data sets from the definition phase through to production deployment (end-to-end).
  • Provide solutions for the design and implementation of Hadoop EMR Cluster/ Big Data Infrastructure.
  • Deploy Hadoop/Big Data/Spark and database storage Infrastructures in AWS cloud.
  • Monitor HDFS/Hadoop/Spark and related software releases, third-party utilities with emphasis on overall system performance.
  • Lead and develop tools and procedures to monitor and automate system tasks on servers and clusters
  • Lead and collaborate with other teams to design, develop, and deploy data tools that support both operations and product use cases.
  • Data Design
  • Lead and design distributed, scalable, and reliable data pipelines that ingest and process data at scale and in real-time.
  • Lead and design big data technologies and prototype solutions to improve data processing architecture.
Qualifications
  • Bachelor’s degree in computer science, computer engineering, or a related technical field.
  • 12+ years of professional experience as a data software engineer; or 16+ years of related experience as a data software engineer in lieu of 4-year degree.
  • 2+ years of experience with AWS cloud or other cloud Big Data computing design, provisioning, and tuning.
  • Related AWS certification, preferred.
  • Previous experience as a Data Engineer / Database Administrator and/or Business Intelligence Analyst.
  • Expertise in database concepts, object and data modeling techniques and design principles.
  • Expertise in database architectures, software, and facilities
  • Expertise with programming languages - Python (required), Scala, Ruby, R Database technologies - SQL, performance tuning concepts, AWS RDS, RedShift, MySQL
  • Expertise with big data batch processing tools: Hadoop MapReduce, ElasticSearch, PIG, Hive, Cascading/Scalding, Apache Spark, AWS EMR
  • Expertise with stream-processing systems: Kinesis, Kafka, MQTT
  • Expertise with relational NoSQL databases including DyanamoDB
  • Expertise in writing JSON, XML, YAML and other data definition schemas
  • Excellent verbal and written communication skills necessary to effectively collaborate in a team environment and present and explain technical information and provide advice to management.
  • Ability to work on advanced complex technical projects or business issues requiring state of the art technical knowledge or industry.
  • Ability to work on significant and unique issues where analysis of situations or data requires an evaluation of intangibles. Exercises independent judgment in methods, techniques and evaluation criteria for obtaining results.
  • Ability to lead and mentor junior engineers and colleagues.

Central Business Solutions, Inc(A Certified Minority Owned Organization)

#J-18808-Ljbffr