1

Databricks Software Jobs in Allentown, PA (NOW HIRING)

Databricks Software information

See Allentown, PA salary details

$47.4K

$110.4K

$163.8K

How much do databricks software jobs pay per year?

As of Aug 8, 2026, the average yearly pay for databricks software in Allentown, PA is $110,372.00, according to ZipRecruiter salary data. Most workers in this role earn between $88,800.00 and $128,300.00 per year, depending on experience, location, and employer.

What is Databricks software?

Databricks Software is a unified analytics platform built on Apache Spark that provides tools for big data processing, machine learning, and collaborative data science. It enables organizations to store, manage, and analyze large datasets efficiently, supporting both batch and streaming data workloads. Databricks also offers collaborative notebooks, automated workflows, and integrations with cloud storage and data lakes, making it a popular choice for data engineering, data science, and business analytics teams.

What are some common challenges faced by Databricks software engineers, and how can they be overcome?

Databricks Software Engineers often encounter challenges related to scaling big data pipelines, optimizing Spark workloads, and integrating diverse data sources. Navigating the complexity of distributed systems and managing cloud infrastructure can be demanding, especially when ensuring data reliability and security. To overcome these challenges, engineers typically collaborate closely with data scientists, DevOps, and platform teams, leverage Databricks' extensive documentation and community support, and adopt best practices such as version control and continuous integration. Regular knowledge sharing and staying updated with new features also help engineers succeed in this dynamic environment.

What are the key skills and qualifications needed to thrive as a Databricks software engineer, and why are they important?

To thrive as a Databricks Software Engineer, you need strong programming skills in languages like Python, Scala, or Java, as well as a solid understanding of distributed computing and data engineering concepts. Familiarity with Databricks platform, Apache Spark, cloud services (such as AWS or Azure), and relevant certifications like Databricks Certified Data Engineer are highly valued. Excellent problem-solving abilities, collaboration, and effective communication are important soft skills for this role. These skills ensure efficient development, deployment, and optimization of big data solutions that drive business insights and innovation.

What is the difference between Databricks Software vs Data Engineer?

AspectDatabricks SoftwareData Engineer
Primary RolePlatform for data analytics and machine learningBuilds, maintains data pipelines and infrastructure
Required SkillsSQL, Spark, cloud platforms, data science basicsSQL, ETL, programming (Python, Scala), database management
Work EnvironmentCloud-based, collaborative data platformData teams, cloud or on-premises environments
CertificationsDatabricks certifications, cloud certificationsNone specific, often cloud or data certifications

While Databricks Software provides a platform for data analytics and machine learning, Data Engineers focus on building and maintaining data pipelines and infrastructure. Both roles often work together but have distinct responsibilities and skill sets within the data ecosystem.

What are popular job titles related to Databricks Software jobs in Allentown, PA? For Databricks Software jobs in Allentown, PA, the most frequently searched job titles are:
What job categories do people searching Databricks Software jobs in Allentown, PA look for? The top searched job categories for Databricks Software jobs in Allentown, PA are:
What cities near Allentown, PA are hiring for Databricks Software jobs? Cities near Allentown, PA with the most Databricks Software job openings:

Data Analytics w/ Data Bricks - HYBRID

André Global, Inc.

Bethlehem, PA • On-site

$113K - $135K/yr

Contractor

Re-posted 4 days ago


Job description

This is a contract-to-hire position with our Insurance client.
  • Must be strong on SQL, Python and Databricks
  • Will undergo 1 hour live coding exercise during the interview

Location: Bethlehem, PA. Resource will be required to work onsite a minimum of 3 days per week in the Bethlehem, PA office. Local candidates preferred. If your candidate is not local to Bethlehem, they will be required to relocate and work onsite from Day 1.
You will:
• Collaborate with data scientists and analysts to understand data requirements and translate them into scalable, high performant data pipeline solutions.
• Support data discovery & data preparation for model development. Perform detailed analysis of raw data sources by applying business context and collaborate with cross-functional teams to transform raw data into curated & certified data assets to be used for ML and BI use cases.
• Collaborate with data science and data engineering team to build scalable and reproducible machine learning pipelines for training and inference.
• Implement machine learning models into operations and processes via batch, streaming and API methods.
• Monitor and troubleshoot data pipeline performance, identifying and resolving bottlenecks and issues.
• Develop, test, and maintain robust tools, frameworks, and libraries that standardize and streamline the data & machine learning lifecycle.
• Contribute to developing and maintaining end-to-end MLOps lifecycle to automate machine learning solutions development and delivery.
• Implement robust monitoring framework for model performance.
• Collaborate with cross-functional teams of Data Science, Data Engineering, business units and various IT teams.
• Create and maintain effective documentation for project and practices ensuring transparency and effective team communication.
You Have:
• Bachelor's or master's degree with 5+ years of experience in Computer Science, Data Science, Engineering, or a related field.
• 4+ years of experience in working with Python, SQL, PySpark and bash scripts. Proficient in software development lifecycle and software engineering practices.
• 2+ years of hands-on experience in using Databricks platform
• 3+ years of hands-on experience in operationalizing Machine Learning solutions which are used in live production processes.
• 2+ years of experience and proficiency in API development using FastAPI frameworks and familiarity with containerization technologies like docker or Kubernetes.
• 3+ years of experience in developing and maintaining robust data pipelines data to be used by Data Scientists to build ML Models.
• 3+ years of experience working with Cloud Data Warehousing (Redshift, Snowflake, Databricks SQL or equivalent) platforms and experience in working with distributed framework like Spark.
• Solid understanding of machine learning life cycle, data mining, and ETL techniques.
• Experience with machine learning frameworks (like Keras or PyTorch) and libraries (like scikit-learn, xgboost).
• Hands-on experience in building and maintaining tools and libraries which have been used by multiple teams across organization.
• Proficient in understanding and incorporating software engineering principles in design & development process.
• Hands on experience with CI/CD tools (e.g., Jenkins or equivalent), version control (Github, Bitbucket), Orchestration (Airflow, Prefect or equivalent)
• Excellent communication skills and ability to work and collaborate with cross functional teams across technology and business.
Good to have:
• Familiarity with deep learning frameworks and deploying deep learning models for production use cases.
• Familiarity in using GPU compute either for model training or inference.
• Understanding of Large language models (LLM) and MLOps lifecycle for operationalizing LLM models.