1

Hadoop Python Jobs in Virginia (NOW HIRING)

Senior Spark & Python Developer

Mclean, VA · On-site

$122K - $165K/yr

Drawing from your vast hands-on experience in Python (PySpark), Spark, REST, Java, and Scala, you ... Experience working in a distributed computing infrastructure like Hadoop and/or Spark * Critical ...

Senior Python Developer with Spark

Reston, VA · On-site

$126K - $170K/yr

Solid experience with big data ecosystems such as Hadoop, Hive, and EMR . Advanced proficiency in ... Python * SQL * Terraform * GitLab What you can expect from us: Together, as owners, let's turn ...

Java, Scala, Python Developer LOCATION Tysons, VA 22182 CLEARANCE TS/SCI Full Poly (Please note ... Familiarity with big data frameworks (e.g., Spark, Hadoop) * Understanding of microservices ...

Java, Scala, Python Developer LOCATION Chantilly, VA 20151 CLEARANCE TS/SCI Full Poly (Please note ... Familiarity with big data frameworks (e.g., Spark, Hadoop) * Understanding of microservices ...

Java, Scala, Python Developer LOCATION Reston, VA 20190 CLEARANCE TS/SCI Full Poly (Please note ... Familiarity with big data frameworks (e.g., Spark, Hadoop) * Understanding of microservices ...

Experience with SQL, Hadoop, Python, Node XL, and a knowledge of telecommunications * Possess a related Bachelor's Degree or equivalent experience. * 8+ years of overall working experience in this ...

ETL Developer

Richmond, VA · On-site

$50.50 - $66/hr

... Hadoop, Spark, Impala, Hive Python. • 3+ years of experience in running, using and troubleshooting the ETL Cloudera Hadoop Ecosystem i.e. Hadoop FS, Hive, Impala, Spark, Kafka, Hue, Oozie, Yarn ...

Experience with SQL, Hadoop, Python, Node XL, and a knowledge of telecommunications * Possess a related Bachelor's Degree or equivalent experience. * 8+ years of overall working experience in this ...

Showing results 21-40

Hadoop Python information

What is a Hadoop Python developer?

A Hadoop Python developer is a software professional who specializes in using Python programming language to develop, implement, and maintain applications that process and analyze large datasets within the Hadoop ecosystem. They leverage Python libraries like PySpark to write scalable data processing scripts, interact with Hadoop components such as HDFS, and optimize big data workflows. These developers play a critical role in building data pipelines, performing data transformation, and supporting analytics projects in organizations that handle vast amounts of data.

What are the key skills and qualifications needed to thrive as a Hadoop Python developer?

To thrive as a Hadoop Python Developer, you need a strong understanding of distributed computing, Hadoop ecosystem components (like HDFS, MapReduce, Hive, or Pig), and advanced Python programming skills, often supported by a degree in computer science or related field. Familiarity with tools such as Apache Spark, Sqoop, and workflow schedulers (like Oozie or Airflow), along with experience in handling big data platforms, is typically required. Problem-solving abilities, attention to detail, and effective communication help developers collaborate with teams and translate business requirements into scalable data solutions. These skills and qualifications are essential for efficiently processing and analyzing large datasets, ensuring data reliability, and driving business insights.

How do Hadoop Python developers typically collaborate with data engineers and analysts on large-scale data projects?

Hadoop Python developers frequently work alongside data engineers and analysts to design, implement, and optimize data pipelines for handling vast datasets. They are responsible for writing Python scripts that interface with Hadoop components, ensuring data is processed efficiently and meets project requirements. Regular communication with data engineers helps align on infrastructure and architectural decisions, while close collaboration with analysts ensures data outputs are accurate and actionable. Agile methodologies and daily stand-ups are common, fostering teamwork and quick problem-solving.

What is the difference between Hadoop Python vs Hadoop Java Developer?

AspectHadoop PythonHadoop Java Developer
Required CredentialsPython programming skills, Hadoop certificationsJava programming skills, Hadoop certifications
Work EnvironmentData analysis, scripting, data pipeline developmentCore development, system integration, big data application coding
Industry UsageData science, analytics, machine learning projectsData infrastructure, platform development, system optimization

Hadoop Python and Hadoop Java Developer roles both involve working with Hadoop ecosystems, but Python focuses more on data analysis and scripting, while Java is geared towards core development and system integration. The choice depends on your programming expertise and career goals within big data environments.

What are jobs in Hadoop Python?

Jobs in Hadoop Python typically involve developing and maintaining data processing tasks using Python scripts within the Hadoop ecosystem. These roles often require knowledge of Hadoop frameworks like MapReduce or Spark, along with Python programming skills, to handle large-scale data analysis and processing tasks. They may also involve working with distributed systems and data pipelines in a big data environment.

What job categories do people searching Hadoop Python jobs in Virginia look for?

The top searched job categories for Hadoop Python jobs in Virginia are:

What cities in Virginia are hiring for Hadoop Python jobs?

Cities in Virginia with the most Hadoop Python job openings:

Senior Spark & Python Developer

Mavens Guild

Mclean, VA • On-site

$122K - $165K/yr

Full-time

Re-posted 3 days ago


Job description

What we would like to see:

In a senior developer role, you will design and build data flow and data integration processes to enhance loss prevention technologies for a leading financing firm. Drawing from your vast hands-on experience in Python (PySpark), Spark, REST, Java, and Scala, you will develop, test, and deploy end-to-end solutions using full-stack development tools within AWS (EMR, S3) cloud based infrastructure.

A typical day as a Senior Spark Programmer:
  • Develop fault tolerant, streaming as well as batch data integration processes using Spark/PySpark, Java and performance enhanced SQL
  • Develop and lead development efforts of application programming interfaces to enable integration of fraud detection systems with a host of new reporting and data mining tools
  • Design and develop automation of data flow tasks and end-to-end process testing
  • Design and develop scalable frameworks to ingest, transform, store, and present loss prevention information to downstream systems
  • Implement and lead implementation of Agile best practices and a continuous integration ecosystem
  • Implement and lead implementation efforts of Spark/Python based solution architecture, scalable process frameworks, advanced analytics, and responsive RESTful services
What you will need to bring to the table:
  • 5+ years experience with processing of structured, unstructured and semi-structured data using in-memory cluster computing technologies, specifically with Spark
  • 5+ years experience as a Java programmer
  • 3+ years experience programming in Python (PySpark API)
  • 3+ years experience with cloud services offered through AWE, like EMR, Redshift, or S3
  • 3+ years experience in developing, testing and deploying RESTful APIs for high volume data streams
  • Experience with continuous integration tools like Jenkins
  • Experience working in a distributed computing infrastructure like Hadoop and/or Spark
  • Critical and analytical approach to solving technical problems
  • Excellent interpersonal skills and ability to clearly communicate highly technical concepts to business stakeholders and technical developers alike