1

Multimodal Learning Jobs (NOW HIRING)

Machine Learning Engineer, Data Mining

Pittsburgh, PA · On-site +1

$111K - $133K/yr

Omnitag, our ML-powered multimodal data mining framework, is the engine that powers this discovery. As a Machine Learning Engineer on the Data Mining team, your mission is to help build the "Brain ...

Machine Learning Engineer, Data Mining

Boston, MA · On-site +1

$124K - $149K/yr

Omnitag, our ML-powered multimodal data mining framework, is the engine that powers this discovery. As a Machine Learning Engineer on the Data Mining team, your mission is to help build the "Brain ...

... multimodal learning, causal ML, generative AI) • Integrate structured and unstructured EHR data with emerging data types such as genomics, imaging, and clinical notes • Partner with clinicians ...

... multimodal learning, causal ML, generative AI) • Integrate structured and unstructured EHR data with emerging data types such as genomics, imaging, and clinical notes • Partner with clinicians ...

Our research spans foundation models, agentic AI, multimodal learning, reasoning systems, scalable training algorithms, evaluation science, inference optimization, and AI systems infrastructure. We ...

next page

Showing results 1-20

Multimodal Learning information

See salary details

$21K

$61.7K

$114.5K

How much do multimodal learning jobs pay per year?

As of Jul 21, 2026, the average yearly pay for multimodal learning in the United States is $61,692.00, according to ZipRecruiter salary data. Most workers in this role earn between $41,000.00 and $72,000.00 per year, depending on experience, location, and employer.

What is multimodal learning?

Multimodal learning is an area of machine learning that involves integrating and processing information from multiple types of data, such as text, images, audio, and video. The goal is to create models that can understand and make predictions based on more than one data modality, similar to how humans use various senses. This approach is used in applications like speech recognition with visual cues, image captioning, and video analysis. By combining different data types, multimodal learning systems can achieve better accuracy and more robust understanding.

What is the difference between Multimodal Learning vs Data Scientist?

AspectMultimodal LearningData Scientist
Required CredentialsAdvanced degrees in AI, Machine Learning, or Computer ScienceBachelor's or Master's in Data Science, Statistics, or related fields
Work EnvironmentResearch labs, AI development teams, academiaBusiness, tech companies, analytics teams
Industry UsageAI research, multimedia applications, roboticsData analysis, predictive modeling, business insights

Multimodal Learning focuses on developing AI models that process and integrate multiple data types like images, text, and audio. Data Scientists analyze data to extract insights, build models, and support decision-making. While both roles involve data and algorithms, Multimodal Learning is specialized in AI model development for complex data integration, whereas Data Scientists work broadly across data analysis and interpretation.

What are the key skills and qualifications needed to thrive as a Multimodal Learning Specialist, and why are they important?

To excel as a Multimodal Learning Specialist, you need a solid background in machine learning, data science, and computer vision, often supported by an advanced degree in a related field. Familiarity with deep learning frameworks like TensorFlow or PyTorch, experience integrating data from diverse sources (e.g., text, audio, images), and knowledge of relevant algorithms are crucial. Strong problem-solving abilities, creativity, and effective collaboration are standout soft skills for this role. These competencies are vital for developing innovative models that can process and interpret complex, multi-source data to drive impactful AI solutions.

What are some common challenges faced by professionals working in multimodal learning roles, and how can they be addressed?

Professionals in multimodal learning frequently encounter challenges related to integrating and aligning data from multiple sources, such as text, images, audio, or video. Ensuring data quality and consistency across modalities can be complex, and developing models that effectively combine heterogeneous information often requires advanced technical skills and innovative thinking. Collaboration with domain experts and other data scientists is key to overcoming these obstacles, as is staying up to date with the latest research and tools in machine learning. Regular team meetings and cross-disciplinary workshops can help foster a collaborative environment and promote knowledge sharing.
More about Multimodal Learning jobs
What cities are hiring for Multimodal Learning jobs? Cities with the most Multimodal Learning job openings:
What states have the most Multimodal Learning jobs? States with the most job openings for Multimodal Learning jobs include:
Infographic showing various Multimodal Learning job openings in the United States as of July 2026, with employment types broken down into 1% As Needed, 72% Full Time, 25% Part Time, 1% Temporary, and 1% Contract. Highlights an 86% Physical, 2% Hybrid, and 12% Remote job distribution, with an average salary of $61,692 per year, or $29.7 per hour.
Machine Learning Engineer, Data Mining

Machine Learning Engineer, Data Mining

Motional

Pittsburgh, PA • On-site, Remote

$111K - $133K/yr

Other

Posted 28 days ago


Job description

Mission Summary:
At Motional, we're transforming how autonomous vehicles discover critical intelligence hidden within petabytes of multimodal sensor data. Our next-generation autonomous driving stack depends on finding the rare edge cases, long-tail scenarios, and model errors that matter most. Omnitag, our ML-powered multimodal data mining framework, is the engine that powers this discovery.
As a Machine Learning Engineer on the Data Mining team, your mission is to help build the "Brain" of this engine. You will work with state-of-the-art foundation models to extract insights from Motional's driving data, working at the intersection of large-scale representation learning and data retrieval. By building smarter mining tools and efficient data pipelines, you will accelerate the model improvement lifecycle for teams working on post-training analysis, error diagnosis, and dataset curation.

What You'll Do:

  • Build and Train ML Pipelines: Develop, train, and fine-tune machine learning models for multimodal sensor data (e.g., vision, LiDAR). Focus on implementing supervised and self-supervised learning approaches to improve data search and retrieval.
  • Support Model Deployment: Implement scalable data preprocessing and augmentation pipelines. Assist in applying standard optimization techniques (e.g., batch inference, quantization) to ensure models run efficiently in production environments.
  • Data Mining & Analysis: Help develop embedding-based search tools and "active learning" workflows to identify critical driving scenarios.
  • Monitor Production Performance: Help build and maintain dashboards to monitor model health, data drift, and system performance. Identify regressions and assist in the operational support of our data mining services.
  • Learn and Apply Best Practices: Follow software engineering standards (version control, CI/CD, unit testing) for ML code. Participate in code reviews and contribute to technical documentation.
  • Collaborate Across Teams: Work closely with senior engineers and machine learning engineers to translate model prototypes into maintainable, scalable engineering solutions.

What We're Looking For (Must-Haves):

  • BS or MS in Computer Science, Machine Learning, or a related field.
  • Hands-on experience with PyTorch (preferred) or TensorFlow/JAX. You should be comfortable training models and evaluating them using standard metrics.
  • Strong proficiency in Python with the ability to write clean, modular, and well-documented code.
  • Working knowledge of version control, unit testing, and basic software design patterns.
  • Experience working with large datasets, including proficiency in SQL and data libraries like Pandas and NumPy.
  • A solid grasp of the full ML lifecycle, from data cleaning and feature engineering to validation and deployment basics.
  • A proactive learner who thrives on constructive feedback and is eager to grow within a high-stakes engineering environment.

Bonus Points (Nice-to-Haves):

  • MS/PhD in Computer Science, Machine Learning, or related field.
  • Experience with agentic systems, autonomous reasoning, chain-of-thought models, or LLM-based planning.
  • Background in autonomous driving, robotics, or real-time decision-making systems.
  • Familiarity with multimodal learning, sensor fusion, or embodied AI.
  • Experience building active learning loops, using the model to find the data that breaks the model.
  • Experience with ML-based data mining, active learning, or contrastive learning.
  • Knowledge of model serving tools (TF Serving, Triton, TorchServe) and MLOps platforms.
  • Publication in top-tier conferences (e.g., ICCV, CVPR, ECCV)

We encourage a hybrid schedule with in-office time at one of our locations in Boston, Pittsburgh, or Las Vegas to support collaboration, or this role can be fully remote.