1

Multimodal Learning Jobs (NOW HIRING)

... multimodal learning, causal ML, generative AI) • Integrate structured and unstructured EHR data with emerging data types such as genomics, imaging, and clinical notes • Partner with clinicians ...

... multimodal learning, causal ML, generative AI) • Integrate structured and unstructured EHR data with emerging data types such as genomics, imaging, and clinical notes • Partner with clinicians ...

Showing results 21-40

Multimodal Learning information

See salary details

$21K

$61.7K

$114.5K

How much do multimodal learning jobs pay per year?

As of Aug 10, 2026, the average yearly pay for multimodal learning in the United States is $61,692.00, according to ZipRecruiter salary data. Most workers in this role earn between $41,000.00 and $72,000.00 per year, depending on experience, location, and employer.

What is multimodal learning?

Multimodal learning is an area of machine learning that involves integrating and processing information from multiple types of data, such as text, images, audio, and video. The goal is to create models that can understand and make predictions based on more than one data modality, similar to how humans use various senses. This approach is used in applications like speech recognition with visual cues, image captioning, and video analysis. By combining different data types, multimodal learning systems can achieve better accuracy and more robust understanding.

What is the difference between Multimodal Learning vs Data Scientist?

AspectMultimodal LearningData Scientist
Required CredentialsAdvanced degrees in AI, Machine Learning, or Computer ScienceBachelor's or Master's in Data Science, Statistics, or related fields
Work EnvironmentResearch labs, AI development teams, academiaBusiness, tech companies, analytics teams
Industry UsageAI research, multimedia applications, roboticsData analysis, predictive modeling, business insights

Multimodal Learning focuses on developing AI models that process and integrate multiple data types like images, text, and audio. Data Scientists analyze data to extract insights, build models, and support decision-making. While both roles involve data and algorithms, Multimodal Learning is specialized in AI model development for complex data integration, whereas Data Scientists work broadly across data analysis and interpretation.

What are the key skills and qualifications needed to thrive in multimodal learning, and why are they important?

To excel as a Multimodal Learning Specialist, you need a solid background in machine learning, data science, and computer vision, often supported by an advanced degree in a related field. Familiarity with deep learning frameworks like TensorFlow or PyTorch, experience integrating data from diverse sources (e.g., text, audio, images), and knowledge of relevant algorithms are crucial. Strong problem-solving abilities, creativity, and effective collaboration are standout soft skills for this role. These competencies are vital for developing innovative models that can process and interpret complex, multi-source data to drive impactful AI solutions.

What are some common challenges faced by professionals working in multimodal learning roles, and how can they be addressed?

Professionals in multimodal learning frequently encounter challenges related to integrating and aligning data from multiple sources, such as text, images, audio, or video. Ensuring data quality and consistency across modalities can be complex, and developing models that effectively combine heterogeneous information often requires advanced technical skills and innovative thinking. Collaboration with domain experts and other data scientists is key to overcoming these obstacles, as is staying up to date with the latest research and tools in machine learning. Regular team meetings and cross-disciplinary workshops can help foster a collaborative environment and promote knowledge sharing.
More about Multimodal Learning jobs
What cities are hiring for Multimodal Learning jobs? Cities with the most Multimodal Learning job openings:
What states have the most Multimodal Learning jobs? States with the most job openings for Multimodal Learning jobs include:
Infographic showing various Multimodal Learning job openings in the United States as of August 2026, with employment types broken down into 1% As Needed, 73% Full Time, 23% Part Time, 1% Temporary, and 2% Contract. Highlights an 88% Physical, 2% Hybrid, and 10% Remote job distribution, with an average salary of $61,692 per year, or $29.7 per hour.

Post Doctoral Researcher - Multimodal Knowledge Extraction and Reasoning

ExxonMobil

Spring, TX • On-site

$106K/yr

Full-time

Medical, Life

Posted 27 days ago


ExxonMobil rating

5.9

Company rating: 5.9 out of 10

Based on 228 frontline employees who took The Breakroom Quiz

71st of 86 rated oil and gas companies


Job description

About us
At ExxonMobil, our vision is to lead in energy innovations that advance modern living while reducing emissions. As one of the world's largest publicly traded energy and chemical companies, we are powered by a unique and diverse workforce fueled by the pride in what we do and what we stand for.
The success of our Upstream, Product Solutions and Low Carbon Solutions businesses is the result of the talent, curiosity and drive of our people. They bring solutions every day to optimize our strategy in energy, chemicals, lubricants and lower-emissions technologies.
We invite you to bring your ideas to ExxonMobil to help create sustainable solutions that improve quality of life and meet society's evolving needs. Learn more about our What and our Why and how we can work together.
About the Role
ExxonMobil is seeking a highly motivated Postdoctoral Researcher specializing in multimodal knowledge extraction and reasoning. The successful candidate will develop advanced AI methods to extract, integrate, and reason over information from diverse data sources-including text, images, video, time series, and structured data-to support critical business and engineering decisions.
This role is ideal for a recent Ph.D. graduate with expertise in multimodal machine learning, knowledge representation, and reasoning systems. The candidate will work in a collaborative environment to build next-generation AI systems that transform complex, heterogeneous data into actionable insights.
Key Responsibilities
  • Develop methods for multimodal data fusion and representation learning across text, visual, spatial, and temporal data.
  • Design models for knowledge extraction, including entity recognition, relation extraction, and structured information generation from unstructured and semi-structured data.
  • Build reasoning systems that combine neural methods with symbolic or knowledge-based approaches.
  • Develop and apply large language model (LLM)-based and multimodal foundation models for knowledge understanding and reasoning.
  • Construct and utilize knowledge graphs and structured representations for enhanced reasoning and decision support.
  • Enable context-aware inference and decision-making using heterogeneous data sources.
  • Evaluate models for accuracy, robustness, and reasoning capability, including explainability where relevant.
  • Collaborate with domain experts to translate extracted knowledge into decision-support workflows.
  • Implement scalable pipelines using modern ML frameworks and data engineering best practices.
  • Communicate findings through technical reports, journal publications, and conference presentations.

Example Research Areas
  • Multimodal machine learning and cross-modal representation learning
  • Knowledge extraction from text, images, and sensor data
  • Knowledge graphs and graph-based reasoning
  • Neural-symbolic AI and hybrid reasoning systems
  • Large language models and multimodal foundation models
  • Information retrieval, semantic search, and question answering
  • Temporal and causal reasoning in complex systems
  • Applications to engineering, scientific, and industrial data environments

Required Qualifications
  • Ph.D. in Computer Science, Artificial Intelligence, Machine Learning, or a closely related field, with a focus on multimodal learning, knowledge extraction, or reasoning.
  • Demonstrated research experience in multimodal machine learning and/or knowledge-based AI, including one or more of:
      • Multimodal representation learning
      • Information extraction or natural language understanding
      • Knowledge graphs or structured representations
      • Reasoning systems (neural, symbolic, or hybrid)
  • Experience with modern deep learning architectures, including transformers and foundation models.
  • Strong programming skills in Python.
  • Hands-on experience with machine learning frameworks such as PyTorch, TensorFlow, or JAX.
  • Experience working with heterogeneous datasets (text, images, structured data, etc.).
  • Strong analytical, problem-solving, and communication skills.
  • Ability to work effectively in multidisciplinary teams.

Preferred Qualifications
  • Experience with multimodal foundation models or large language models (LLMs).
  • Familiarity with knowledge graph construction, querying, and reasoning frameworks.
  • Experience with retrieval-augmented generation (RAG) or hybrid search systems.
  • Background in probabilistic reasoning, causal inference, or uncertainty-aware AI.
  • Experience with scalable data pipelines and distributed ML systems.
  • Experience applying AI methods to scientific, engineering, or industrial datasets.
  • Strong publication record in multimodal AI, NLP, or knowledge-based systems.
  • Demonstrated ability to translate research into practical decision-support tools.

Duration
This opportunity is for a postdoctoral position expected to last one to three years, subject to annual review and renewal.
Work Location
Our post doctoral research employees are located at our main corporate office in Spring, Texas.
Your Total Rewards
An ExxonMobil career is one designed to last. Our commitment to you runs deep: our employees grow personally and professionally, with benefits built on our core categories of health, security, finance, and life. Individual pay is determined based on various factors including degree/education, discipline, year of study, skills, abilities, qualifications, and work experience.
More information on our Company's benefits can be found at www.exxonmobilfamily.com.
Please note pay rates and benefits may be changed from time to time without notice, subject to applicable law.
Relocation Options
Relocation benefits may be available to you based on ExxonMobil eligibility guidelines.
Equal Opportunity Employer
ExxonMobil is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, age, sexual orientation, gender identity, national origin, citizenship status, protected veteran status, genetic information, or physical or mental disability.
Nothing herein is intended to override the corporate separateness of local entities. Working relationships discussed herein do not necessarily represent a reporting connection, but may reflect a functional guidance, stewardship, or service relationship.
Exxon Mobil Corporation has numerous affiliates, many with names that include ExxonMobil, Exxon, Esso and Mobil. For convenience and simplicity, those terms and terms like corporation, company, our, we and its are sometimes used as abbreviated references to specific affiliates or affiliate groups. Abbreviated references describing global or regional operational organizations and global or regional business lines are also sometimes used for convenience and simplicity. Similarly, ExxonMobil has business relationships with thousands of customers, suppliers, governments, and others. For convenience and simplicity, words like venture, joint venture, partnership, co-venturer, and partner are used to indicate business relationships involving common activities and interests, and those words may not indicate precise legal relationships.

What ExxonMobil employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom