Focus Multimodal Foundation Models Representation Learning Method Innovation We are looking for strong technical builders and researchers who deeply understand foundation models and representation ...
Focus Multimodal Foundation Models Representation Learning Method Innovation We are looking for strong technical builders and researchers who deeply understand foundation models and representation ...
Focus Multimodal Foundation Models · Representation Learning · Method Innovation We are looking for strong technical builders and researchers who deeply understand foundation models and ...
Quick apply
Focus Multimodal Foundation Models · Representation Learning · Method Innovation We are looking for strong technical builders and researchers who deeply understand foundation models and ...
Focus Multimodal Foundation Models • Representation Learning • Method Innovation We are looking for strong technical builders and researchers who deeply understand foundation models and ...
Focus Multimodal Foundation Models • Representation Learning • Method Innovation We are looking for strong technical builders and researchers who deeply understand foundation models and ...
Machine Learning Engineer, Multimodal Perception and Authentication
San Francisco, CA · On-site
$150 - $200/hr
About the Role We're looking for a machine learning engineer to help shape how future AI systems understand the physical world and the people in it. The role focuses on multimodal perception and ...
Machine Learning Engineer, Multimodal Perception and Authentication
San Francisco, CA · On-site
$150 - $200/hr
About the Role We're looking for a machine learning engineer to help shape how future AI systems understand the physical world and the people in it. The role focuses on multimodal perception and ...
$150 - $200/hr
About the Role We're looking for a machine learning engineer to help shape how future AI systems understand the physical world and the people in it. The role focuses on multimodal perception and ...
$150 - $200/hr
About the Role We're looking for a machine learning engineer to help shape how future AI systems understand the physical world and the people in it. The role focuses on multimodal perception and ...
Machine Learning Engineer - Multimodal Modeling
San Francisco, CA · On-site
$250K - $295K/yr
As a Machine Learning Engineer on the Applied Science team, you will design, train, and deploy Stand's flagship AI capabilities, with a central focus on the multimodal meshing of our Stand World ...
Machine Learning Engineer - Multimodal Modeling
San Francisco, CA · On-site
$250K - $295K/yr
As a Machine Learning Engineer on the Applied Science team, you will design, train, and deploy Stand's flagship AI capabilities, with a central focus on the multimodal meshing of our Stand World ...
Machine Learning Engineer, Multimodal Perception and Authentication
San Francisco, CA · On-site
$150 - $230/hr
About the Role We're looking for a machine learning engineer to help shape how future AI systems understand the physical world and the people in it. The role focuses on multimodal perception and ...
Machine Learning Engineer, Multimodal Perception and Authentication
San Francisco, CA · On-site
$150 - $230/hr
About the Role We're looking for a machine learning engineer to help shape how future AI systems understand the physical world and the people in it. The role focuses on multimodal perception and ...
Machine Learning Engineer, Multimodal Perception and Authentication
San Francisco, CA · On-site
$342K - $399K/yr
About the Role We're looking for a machine learning engineer to help shape how future AI systems understand the physical world and the people in it. The role focuses on multimodal perception and ...
Machine Learning Engineer, Multimodal Perception and Authentication
San Francisco, CA · On-site
$342K - $399K/yr
About the Role We're looking for a machine learning engineer to help shape how future AI systems understand the physical world and the people in it. The role focuses on multimodal perception and ...
AI Scientist - Biomedical Multimodal Modeling
South San Francisco, CA · On-site
$170K - $240K/yr
... multimodal data ... Our approach combines representation learning and generative modeling to capture structure ...
AI Scientist - Biomedical Multimodal Modeling
South San Francisco, CA · On-site
$170K - $240K/yr
... multimodal data ... Our approach combines representation learning and generative modeling to capture structure ...
Help train and develop multimodal learning models using advanced learning techniques including RAG, self-supervised learning, semi-supervised, and transductive learning. Requirements Desired ...
Help train and develop multimodal learning models using advanced learning techniques including RAG, self-supervised learning, semi-supervised, and transductive learning. Requirements Desired ...
Help train and develop multimodal learning models using advanced learning techniques including RAG, self-supervised learning, semi-supervised, and transductive learning. Requirements Desired ...
Help train and develop multimodal learning models using advanced learning techniques including RAG, self-supervised learning, semi-supervised, and transductive learning. Requirements Desired ...
Member of Technical Staff, Multimodal Vision
$180K - $450K/yr
Strong experience in image understanding, video modeling, or multimodal learning. * Strong background in data-driven experimentation, evaluation design, and iterative model development. Bonus ...
Member of Technical Staff, Multimodal Vision
$180K - $450K/yr
Strong experience in image understanding, video modeling, or multimodal learning. * Strong background in data-driven experimentation, evaluation design, and iterative model development. Bonus ...
Senior Staff Machine Learning Scientist, Assets
OR · On-site +1
Design, implement, train, and optimize large-scale vision and multimodal foundation models across ... Proficiency in modern deep learning frameworks such as PyTorch and TensorFlow. * Demonstrated ...
Senior Staff Machine Learning Scientist, Assets
OR · On-site +1
Design, implement, train, and optimize large-scale vision and multimodal foundation models across ... Proficiency in modern deep learning frameworks such as PyTorch and TensorFlow. * Demonstrated ...
Help train and develop multimodal learning models using advanced learning techniques including RAG, self-supervised learning, semi-supervised, and transductive learning. Requirements Desired ...
Quick apply
Help train and develop multimodal learning models using advanced learning techniques including RAG, self-supervised learning, semi-supervised, and transductive learning. Requirements Desired ...
AI Scientist - Biomedical Multimodal Modeling
South San Francisco, CA · On-site
$170 - $240/hr
... multimodal data ... Our approach combines representation learning and generative modeling to capture structure ...
AI Scientist - Biomedical Multimodal Modeling
South San Francisco, CA · On-site
$170 - $240/hr
... multimodal data ... Our approach combines representation learning and generative modeling to capture structure ...
Member of Technical Staff, Multimodal Vision
San Jose, CA · On-site
$180K - $450K/yr
Strong experience in image understanding, video modeling, or multimodal learning. * Strong background in data-driven experimentation, evaluation design, and iterative model development. Bonus ...
Member of Technical Staff, Multimodal Vision
San Jose, CA · On-site
$180K - $450K/yr
Strong experience in image understanding, video modeling, or multimodal learning. * Strong background in data-driven experimentation, evaluation design, and iterative model development. Bonus ...
Design, implement, train, and optimize large-scale vision and multimodal foundation models across ... Proficiency in modern deep learning frameworks such as PyTorch and TensorFlow. * Demonstrated ...
Design, implement, train, and optimize large-scale vision and multimodal foundation models across ... Proficiency in modern deep learning frameworks such as PyTorch and TensorFlow. * Demonstrated ...
Machine Learning Scientist
San Francisco, CA · On-site
$180 - $270/hr
As a Machine Learning Scientist, you will develop cutting‑edge AI models to integrate and decode complex, multimodal data streams from our custom sensing hardware. You'll play a pivotal role in ...
Machine Learning Scientist
San Francisco, CA · On-site
$180 - $270/hr
As a Machine Learning Scientist, you will develop cutting‑edge AI models to integrate and decode complex, multimodal data streams from our custom sensing hardware. You'll play a pivotal role in ...
Post Doctoral Researcher - Multimodal Knowledge Extraction and Reasoning
Spring, TX · On-site +1
$106K/yr
D. graduate with expertise in multimodal machine learning, knowledge representation, and reasoning systems. The candidate will work in a collaborative environment to build next-generation AI systems ...
Post Doctoral Researcher - Multimodal Knowledge Extraction and Reasoning
Spring, TX · On-site +1
$106K/yr
D. graduate with expertise in multimodal machine learning, knowledge representation, and reasoning systems. The candidate will work in a collaborative environment to build next-generation AI systems ...
Machine Learning Scientist
San Francisco, CA · On-site
$180K - $270K/yr
Develop and evaluate multimodal learning techniques to fuse information from multiple sensor modalities. * Iterate rapidly on model prototypes for real-time inference on custom hardware. * Create and ...
Machine Learning Scientist
San Francisco, CA · On-site
$180K - $270K/yr
Develop and evaluate multimodal learning techniques to fuse information from multiple sensor modalities. * Iterate rapidly on model prototypes for real-time inference on custom hardware. * Create and ...
Multimodal Learning information
See salary details
$21K - $29.5K
10% of jobs
$29.5K - $38K
14% of jobs
$39.2K is the 25th percentile. Wages below this are outliers.
$38K - $46.5K
10% of jobs
$46.5K - $55K
12% of jobs
The median wage is $57K / yr.
$55K - $63.5K
20% of jobs
$68.8K is the 75th percentile. Wages above this are outliers.
$63.5K - $72K
15% of jobs
$72K - $80.5K
4% of jobs
$80.5K - $89K
2% of jobs
$89K - $97.5K
4% of jobs
$97.5K - $106K
0% of jobs
$106K - $114.5K
9% of jobs
$21K
$61.7K
$114.5K
How much do multimodal learning jobs pay per year?
What is multimodal learning?
What are the key skills and qualifications needed to thrive in multimodal learning, and why are they important?
What are some common challenges faced by professionals working in multimodal learning roles, and how can they be addressed?
What is the difference between Multimodal Learning vs Data Scientist?
| Aspect | Multimodal Learning | Data Scientist |
|---|---|---|
| Required Credentials | Advanced degrees in AI, Machine Learning, or Computer Science | Bachelor's or Master's in Data Science, Statistics, or related fields |
| Work Environment | Research labs, AI development teams, academia | Business, tech companies, analytics teams |
| Industry Usage | AI research, multimedia applications, robotics | Data analysis, predictive modeling, business insights |
Multimodal Learning focuses on developing AI models that process and integrate multiple data types like images, text, and audio. Data Scientists analyze data to extract insights, build models, and support decision-making. While both roles involve data and algorithms, Multimodal Learning is specialized in AI model development for complex data integration, whereas Data Scientists work broadly across data analysis and interpretation.
What cities are hiring for Multimodal Learning jobs?
Cities with the most Multimodal Learning job openings:
What states have the most Multimodal Learning jobs?
States with the most job openings for Multimodal Learning jobs include:
What job categories do people searching Multimodal Learning jobs look for?
The top searched job categories for Multimodal Learning jobs are:

Member of Technical Staff (MTS) - Multimodal Foundation Models
Fremont, CA
Full-time
Re-posted 5 days ago
Job description
Focus
Multimodal Foundation Models Representation Learning Method Innovation
We are looking for strong technical builders and researchers who deeply understand foundation models and representation learning beyond simply applying existing frameworks.
Ideal candidates should have:
- Strong experimental rigor
- Solid systems and modeling intuition
- Hands-on engineering ability
- Interest in scalable multimodal AI systems for real-world autonomy
We value people who can bridge research and production, and who care about robustness, scalability, efficiency, and practical deployment in large-scale autonomous driving systems.
Responsibilities
1. Large-Scale Foundation Model Pretraining
- Develop scalable pretraining pipelines for large-scale multimodal driving data
- Design and optimize training strategies for:
- Vision-language-action models
- Video foundation models
- Long-context temporal modeling
- Multimodal representation alignment
- Improve:
- Training stability
- Data efficiency
- Scaling efficiency
- Representation robustness
- Work on distributed training systems and large-scale model optimization using frameworks such as:
- PyTorch Distributed
- DeepSpeed
- Megatron-LM
2. Representation Learning & Method Innovation
- Design and improve self-supervised and multimodal learning methods for real-world autonomous driving systems
- Conduct architecture-level research on:
- Vision Transformers (ViT)
- Video / temporal architectures
- Multimodal fusion and alignment
- Embedding and retrieval systems
- Long-context and memory-efficient architectures
- Explore and improve:
- Pretraining objectives
- Loss functions
- Training paradigms
- Generalization and robustness
- Analyze model behavior through:
- Rigorous ablation studies
- Failure case analysis
- Representation probing and evaluation
3. Efficient Foundation Models & Scalable Deployment
- Improve the efficiency, scalability, and deployability of large multimodal foundation models for real-world autonomous driving systems
- Work on areas such as:
- Model quantization
- Knowledge distillation
- Efficient attention mechanisms
- Sparse architectures and Mixture-of-Experts (MoE)
- Long-context and memory-efficient modeling
- Inference acceleration and serving optimization
- Training and inference system efficiency
- Optimize model throughput, latency, memory usage, and deployment performance for large-scale production environments
Requirements
- MS or PhD in:
- Computer Vision
- Machine Learning
- Robotics
- Computer Science
- Related fields
- Strong understanding of:
- Foundation models
- Self-supervised learning
- Representation learning
- Multimodal learning
- Large-scale pretraining
- Hands-on experience with methods such as:
- CLIP
- DINO / DINOv2
- MAE
- Contrastive learning
- Masked modeling
- MoE or scalable transformer architectures
- Experience with one or more of the following is highly valued:
- Video foundation models
- Long-context modeling
- Retrieval systems
- Efficient inference
- Distributed training
- Model compression and deployment optimization
- Strong publication record in top-tier venues is preferred:
- CVPR
- ICCV
- ECCV
- NeurIPS
- ICLR
- ICML