1

Hadoop Python Jobs in Virginia (NOW HIRING)

ETL Developer

Richmond, VA · On-site

$50.50 - $66/hr

... Hadoop, Spark, Impala, Hive Python. • 3+ years of experience in running, using and troubleshooting the ETL Cloudera Hadoop Ecosystem i.e. Hadoop FS, Hive, Impala, Spark, Kafka, Hue, Oozie, Yarn ...

Experience with SQL, Hadoop, Python, Node XL, and a knowledge of telecommunications * Possess a related Bachelor's Degree or equivalent experience. * 8+ years of overall working experience in this ...

Senior Software Developer, Data Analytics

Mclean, VA · On-site

$55 - $72.75/hr

Java, Python, R, C#, C, SAS, analytic engines, Hadoop, parallelized analytic algorithms, and NoSQL and massively parallel processing databases. The successful candidate will have the ability to ...

Strong programming skills in Java, Scala, or Python * Familiarity with big data processing tools and techniques * Experience with the Hadoop ecosystem * Good understanding of distributed systems

Experience developing algorithms with Python, SQL, or NoSQL * Experience with MapReduce, Hadoop, Hive, EMR, Spark, Gurobi, or MySQL * Security+ Certification is a nice-to-have Core Competencies ...

New

Senior Data Engineer

Arlington, VA · On-site

$121K - $165K/yr

Spark (advance/expert), Hadoop platform & tools (Hive, Impala, Nifi, Oozie, Scoop) * Databricks * Python, SQL Ideal Candidate Qualifications: * 5+ years of full stack engineering experience in an ...

Java, MapReduce, Python, Pig Programming, Hadoop Streaming, HiveQL * Experience developing RESTful Web Services * Agile/scrum experience * Experience leading and managing large scale, complex ...

Big Data Engineer

Reston, VA · On-site

$58 - $76.75/hr

Java, MapReduce, Python, Pig Programming, Hadoop Streaming, HiveQL * Experience developing RESTful Web Services * Agile/scrum experience * Experience leading and managing large scale, complex ...

Showing results 41-60

Hadoop Python information

What is a Hadoop Python developer?

A Hadoop Python developer is a software professional who specializes in using Python programming language to develop, implement, and maintain applications that process and analyze large datasets within the Hadoop ecosystem. They leverage Python libraries like PySpark to write scalable data processing scripts, interact with Hadoop components such as HDFS, and optimize big data workflows. These developers play a critical role in building data pipelines, performing data transformation, and supporting analytics projects in organizations that handle vast amounts of data.

What are the key skills and qualifications needed to thrive as a Hadoop Python developer?

To thrive as a Hadoop Python Developer, you need a strong understanding of distributed computing, Hadoop ecosystem components (like HDFS, MapReduce, Hive, or Pig), and advanced Python programming skills, often supported by a degree in computer science or related field. Familiarity with tools such as Apache Spark, Sqoop, and workflow schedulers (like Oozie or Airflow), along with experience in handling big data platforms, is typically required. Problem-solving abilities, attention to detail, and effective communication help developers collaborate with teams and translate business requirements into scalable data solutions. These skills and qualifications are essential for efficiently processing and analyzing large datasets, ensuring data reliability, and driving business insights.

How do Hadoop Python developers typically collaborate with data engineers and analysts on large-scale data projects?

Hadoop Python developers frequently work alongside data engineers and analysts to design, implement, and optimize data pipelines for handling vast datasets. They are responsible for writing Python scripts that interface with Hadoop components, ensuring data is processed efficiently and meets project requirements. Regular communication with data engineers helps align on infrastructure and architectural decisions, while close collaboration with analysts ensures data outputs are accurate and actionable. Agile methodologies and daily stand-ups are common, fostering teamwork and quick problem-solving.

What is the difference between Hadoop Python vs Hadoop Java Developer?

AspectHadoop PythonHadoop Java Developer
Required CredentialsPython programming skills, Hadoop certificationsJava programming skills, Hadoop certifications
Work EnvironmentData analysis, scripting, data pipeline developmentCore development, system integration, big data application coding
Industry UsageData science, analytics, machine learning projectsData infrastructure, platform development, system optimization

Hadoop Python and Hadoop Java Developer roles both involve working with Hadoop ecosystems, but Python focuses more on data analysis and scripting, while Java is geared towards core development and system integration. The choice depends on your programming expertise and career goals within big data environments.

What are jobs in Hadoop Python?

Jobs in Hadoop Python typically involve developing and maintaining data processing tasks using Python scripts within the Hadoop ecosystem. These roles often require knowledge of Hadoop frameworks like MapReduce or Spark, along with Python programming skills, to handle large-scale data analysis and processing tasks. They may also involve working with distributed systems and data pipelines in a big data environment.

What job categories do people searching Hadoop Python jobs in Virginia look for?

The top searched job categories for Hadoop Python jobs in Virginia are:

What cities in Virginia are hiring for Hadoop Python jobs?

Cities in Virginia with the most Hadoop Python job openings:

ETL Developer

Richmond, VA • On-site

Apex Informatics
IT Services • 1 - 10 employees

$50.50 - $66/hr

Contractor

Re-posted 9 days ago


Job description

ETL Developer
ATLANTA, GA
RICHMOND VA
Hybrid Role (Look for candidates within GA, VA)
• Strong working knowledge of ETL, database technologies, big data and data processing skills
• 3+ years of experience developing ETL solutions using any ETL tool like Informatica, SSIS, etc.,
• 3+ years of experience developing applications using Hadoop, Spark, Impala, Hive Python.
• 3+ years of experience in running, using and troubleshooting the ETL Cloudera Hadoop Ecosystem i.e. Hadoop FS, Hive, Impala, Spark, Kafka, Hue, Oozie, Yarn, Sqoop, Flume.
• Experience on Autosys JIL scripting.
• Proficient scripting skills i.e. Unix shell Perl Scripting
• Experience troubleshooting data-related issues.
• Experience processing large amounts of structured and unstructured data with MapReduce.
• Experience in SQL and Relation Database developing data extraction applications.
• Experience with data movement and transformation technologies.
• Good to have experience in Python / Scala programming