1

Data Preprocessing Jobs in Texas (NOW HIRING)

Develop and optimize data processing pipelines for data preprocessing, feature engineering, model training, and evaluation. * Data Integration: Evaluate and integrate new data sources to continuously ...

... data analysis, including preprocessing, feature engineering, and leveraging Generative AI algorithms for novel solutions. Β· Lead cross-functional collaborations to integrate Generative AI models ...

... data cleaning and preprocessing β€’ Bachelor's Degree in Computer Science, Data Science, Statistics, Mathematics, Applied Mathematics, Engineering, Economics, Physics, Operations Research ...

... data analysis, including preprocessing, feature engineering, and leveraging Generative AI algorithms for novel solutions. β€’ Lead cross-functional collaborations to integrate Generative AI models ...

Data Scientist

Austin, TX Β· On-site +1

Perform preprocessing of structured and unstructured data * Design, implement and deliver maintainable and high-quality code using best practices (e.g. Git/Github, Secrets, Configurations, Yaml/JSON)

Data Scientist

Austin, TX Β· On-site +1

Perform preprocessing of structured and unstructured data * Design, implement and deliver maintainable and high-quality code using best practices (e.g. Git/Github, Secrets, Configurations, Yaml/JSON)

Services of the platform include data preprocessing, AI model fine-tuning, inference oversight and MLOps for AI solutions in production. For traditional enterprises that have an aspiration to ...

Guides students through data preprocessing, feature selection, building and comparing classification and regression models, implementing clustering algorithms, and interpreting confusion matrices and ...

Showing results 41-60

Data Preprocessing information

What is data preprocessing?

Data preprocessing is the process of cleaning, transforming, and organizing raw data into a usable format for analysis or machine learning. It involves steps such as handling missing values, removing duplicates, normalizing or scaling data, and encoding categorical variables. Proper data preprocessing helps improve the quality and performance of predictive models by ensuring the data is accurate, consistent, and suitable for analysis.

What are the key skills and qualifications needed to thrive as a data preprocessing specialist, and why are they important?

To thrive as a Data Preprocessing Specialist, you need a strong background in statistics, data cleaning, and data transformation, often supported by a degree in computer science, data science, or a related field. Proficiency with tools such as Python (pandas, NumPy), SQL, and data visualization platforms is typically essential, along with familiarity with data management systems. Attention to detail, problem-solving abilities, and effective communication are standout soft skills in this position. These skills are crucial for ensuring high-quality, reliable datasets that underpin accurate data analysis and machine learning outcomes.

What are some common challenges faced in a data preprocessing role, and how can they be effectively managed?

Professionals in Data Preprocessing often encounter challenges such as handling incomplete or inconsistent data, managing large datasets, and ensuring data quality before analysis. Addressing these issues typically involves using specialized tools to automate data cleaning, establishing clear data validation rules, and collaborating closely with data engineers and analysts. Staying updated with best practices and leveraging scripting languages like Python or R can also streamline the preprocessing workflow, making it easier to deliver reliable and accurate datasets for downstream analysis.

What is the difference between Data Preprocessing vs Data Analysis?

AspectData PreprocessingData Analysis
Primary FocusCleaning, transforming, and preparing raw data for analysisInterpreting data to extract insights and support decision-making
Skills RequiredData cleaning, scripting, understanding of data formatsStatistical analysis, data visualization, critical thinking
Work EnvironmentData engineering teams, data science projectsBusiness intelligence, research, data science teams
Tools UsedPython, R, SQL, ETL toolsExcel, Tableau, R, Python, statistical software

While data preprocessing involves preparing raw data for analysis by cleaning and transforming it, data analysis focuses on interpreting the prepared data to uncover trends and insights. Both roles are essential in the data pipeline but serve different purposes in the data lifecycle.

What cities in Texas are hiring for Data Preprocessing jobs?

Cities in Texas with the most Data Preprocessing job openings:

Infographic showing various Data Preprocessing job openings in Texas as of June 2026, with employment types broken down into 40% Internship, and 60% Full Time. Highlights an 100% In-person job distribution.

Sr. Engineer - Machine Learning | Dallas, TX(Day One Onsite)

Dallas, TX β€’ On-site

$121K - $160K/yr

Other

This job post hasΒ expired 1 day ago.Β Applications are no longer accepted.


Job description

Sr. Engineer – Machine Learning

Dallas, TX (Day One Onsite)

Contract

Description:

Collaborate with leaders, business analysts, project managers, IT architects, technical leads and other engineers, along with internal customers, to understand requirements and develop needs according to business requirements for AI solutions

Maintain and enhance existing enterprise services, applications, and platforms using domain driven design and test-driven development

Troubleshoot and debug complex issues; identify and implement solutions

Create detailed project specifications, requirements, and estimates

Research and implement new AI technologies to enhance current processes, security, and performance

Work closely with data scientists and product teams to build and deploy machine learning models, focusing on the technical aspects of model deployment.

Implement and optimize Python-based ML pipelines for data preprocessing, model training, and deployment.

Monitor model performance and implement strategies for bias mitigation and explainability.

Responsible for ensuring models are scalable and efficient in production environments.

Write and maintain code for model training and deployment, collaborating with software engineers to integrate models into applications.

Partner with a diverse team of experts, leveraging cutting-edge technologies to build scalable and impactful AI solutions.

Minimum Qualifications – Education & Prior Job Experience:

Bachelor’s degree in computer science, Computer Engineering, Data Science, Information Systems (CIS/MIS), Engineering or related technical discipline, or equivalent experience/training

7 to 9+ years of full Software Development Life Cycle (SDLC) experience designing, developing, and implementing large-scale machine learning applications in hosted production environments

2+ years of professional, design, and open-source experience

Top Requirements:
  • Mixed-Integer Programming (MIP)modeling; formulatingreal-world businessrules asdecision variables, linearconstraints, andobjectives; LP/IP theory; branch-and-bound intuition.
  • Commercial solverexperience; FICO Xpress (stronglypreferred) or Gurobi/CPLEX/OR-Tools/Hexaly, including solvertuning (gaps, threads, seeds, determinism) andreading solver logs/.lp files.
  • Programming to implement model & visualizemodel solutions; Java (themodellives inJava behind a solverabstraction); Python (incl. Streamlit for thedashboard); ableto code and ship changes in both programming languages.
  • Also required: comfortableusing AI coding agents(Claude/Copilot) for fastchanges; high codingstandards; a collaborative teamplayer (open toideas/feedback, no solo work);a fastlearner.
  • MSor PhD inOperations Research/ IndustrialEngineering /Applied Math (or related), plus3+years appliedoptimization or equivalentdemonstrated MIP experience.
Education:

MS/PhD in OR /IE /Applied Math /CS-optimizationor equivalent.

Nice To Have:
  • Java21 / SpringBoot familiarity
  • clean/hexagonal model design (ports & adapters)
  • TDD (JUnit5) foroptimizationmodels(LP-file integration tests)
  • scheduling/assignment/routing problem experience
  • heuristics/metaheuristics/CP as complements toMIP
  • dataanalysis & debugging
  • mining datasets (incl. AzureADLS) andsolver logs to findissues and derive insights
  • data wrangling (CSV/Excel/SQL)
  • Git/GitHub, MongoDB Compass, Maven,Docker
  • aviation / MRO dom
#J-18808-Ljbffr