Focus Multimodal Foundation Models Representation Learning Method Innovation We are looking for strong technical builders and researchers who deeply understand foundation models and representation ...
Focus Multimodal Foundation Models Representation Learning Method Innovation We are looking for strong technical builders and researchers who deeply understand foundation models and representation ...
Focus Multimodal Foundation Models · Representation Learning · Method Innovation We are looking for strong technical builders and researchers who deeply understand foundation models and ...
Quick apply
Focus Multimodal Foundation Models · Representation Learning · Method Innovation We are looking for strong technical builders and researchers who deeply understand foundation models and ...
Focus Multimodal Foundation Models • Representation Learning • Method Innovation We are looking for strong technical builders and researchers who deeply understand foundation models and ...
Focus Multimodal Foundation Models • Representation Learning • Method Innovation We are looking for strong technical builders and researchers who deeply understand foundation models and ...
Machine Learning Engineer - Multimodal Modeling
San Francisco, CA · On-site
$250K - $295K/yr
As a Machine Learning Engineer on the Applied Science team, you will design, train, and deploy Stand's flagship AI capabilities, with a central focus on the multimodal meshing of our Stand World ...
Machine Learning Engineer - Multimodal Modeling
San Francisco, CA · On-site
$250K - $295K/yr
As a Machine Learning Engineer on the Applied Science team, you will design, train, and deploy Stand's flagship AI capabilities, with a central focus on the multimodal meshing of our Stand World ...
Research Scientist in Multimodal Interaction and World Model - Seed - Graduates - 2027 Start (PhD)
San Jose, CA · On-site
... learning. • Improve agent capabilities such as perception, memory, decision-making, and tool use ... etc. • Experience in multimodal learning, reinforcement learning, or agent systems. • ...
Research Scientist in Multimodal Interaction and World Model - Seed - Graduates - 2027 Start (PhD)
San Jose, CA · On-site
... learning. • Improve agent capabilities such as perception, memory, decision-making, and tool use ... etc. • Experience in multimodal learning, reinforcement learning, or agent systems. • ...
... multimodal data ... Our approach combines representation learning and generative modeling to capture structure ...
... multimodal data ... Our approach combines representation learning and generative modeling to capture structure ...
Help train and develop multimodal learning models using advanced learning techniques including RAG, self-supervised learning, semi-supervised, and transductive learning. Requirements Desired ...
Help train and develop multimodal learning models using advanced learning techniques including RAG, self-supervised learning, semi-supervised, and transductive learning. Requirements Desired ...
Help train and develop multimodal learning models using advanced learning techniques including RAG, self-supervised learning, semi-supervised, and transductive learning. Requirements Desired ...
Help train and develop multimodal learning models using advanced learning techniques including RAG, self-supervised learning, semi-supervised, and transductive learning. Requirements Desired ...
Help train and develop multimodal learning models using advanced learning techniques including RAG, self-supervised learning, semi-supervised, and transductive learning. Requirements Desired ...
Quick apply
Help train and develop multimodal learning models using advanced learning techniques including RAG, self-supervised learning, semi-supervised, and transductive learning. Requirements Desired ...
Member of Technical Staff, Multimodal Vision
$180K - $450K/yr
Experience with large-scale machine learning systems and distributed training. * Strong background ... Experience with multimodal systems (vision + text, vision + audio) or real-time AI systems is a ...
Member of Technical Staff, Multimodal Vision
$180K - $450K/yr
Experience with large-scale machine learning systems and distributed training. * Strong background ... Experience with multimodal systems (vision + text, vision + audio) or real-time AI systems is a ...
Member of Technical Staff, Multimodal Vision
San Jose, CA · On-site
$180K - $450K/yr
Experience with large-scale machine learning systems and distributed training. * Strong background ... Experience with multimodal systems (vision + text, vision + audio) or real-time AI systems is a ...
Member of Technical Staff, Multimodal Vision
San Jose, CA · On-site
$180K - $450K/yr
Experience with large-scale machine learning systems and distributed training. * Strong background ... Experience with multimodal systems (vision + text, vision + audio) or real-time AI systems is a ...
Machine Learning Scientist
San Francisco, CA · On-site
$180K - $270K/yr
As a Machine Learning Scientist, you will develop cutting-edge AI models to integrate and decode complex, multimodal data streams from our custom sensing hardware. You'll play a pivotal role in ...
Machine Learning Scientist
San Francisco, CA · On-site
$180K - $270K/yr
As a Machine Learning Scientist, you will develop cutting-edge AI models to integrate and decode complex, multimodal data streams from our custom sensing hardware. You'll play a pivotal role in ...
Senior Applied AI Researcher, Digital Biology
Santa Clara, CA · On-site
$102K - $125K/yr
The role involves conceptualizing and implementing deep learning architectures for biological data, developing multimodal learning systems, and collaborating with a diverse team to advance ...
Senior Applied AI Researcher, Digital Biology
Santa Clara, CA · On-site
$102K - $125K/yr
The role involves conceptualizing and implementing deep learning architectures for biological data, developing multimodal learning systems, and collaborating with a diverse team to advance ...
Description As a Machine Learning Research Engineer, you will help design and develop models and algorithms for multimodal perception and reasoning leveraging Vision-Language Models (VLMs) and ...
Description As a Machine Learning Research Engineer, you will help design and develop models and algorithms for multimodal perception and reasoning leveraging Vision-Language Models (VLMs) and ...
Research expertise in video generation/understanding, multimodal learning, or diffusion models * Demonstrated significant industry influence in the field of AI and/or recently published research in ...
Research expertise in video generation/understanding, multimodal learning, or diffusion models * Demonstrated significant industry influence in the field of AI and/or recently published research in ...
Machine Learning Research Engineer, SIML - ISE
$150K - $277K/yr
Description As a Machine Learning Research Engineer, you will help design and develop models and algorithms for multimodal perception and reasoning leveraging Vision-Language Models (VLMs) and ...
Machine Learning Research Engineer, SIML - ISE
$150K - $277K/yr
Description As a Machine Learning Research Engineer, you will help design and develop models and algorithms for multimodal perception and reasoning leveraging Vision-Language Models (VLMs) and ...
Senior AI Researcher
San Francisco, CA · On-site
Our research spans foundation models, agentic AI, multimodal learning, reasoning systems, scalable training algorithms, evaluation science, inference optimization, and AI systems infrastructure. We ...
Senior AI Researcher
San Francisco, CA · On-site
Our research spans foundation models, agentic AI, multimodal learning, reasoning systems, scalable training algorithms, evaluation science, inference optimization, and AI systems infrastructure. We ...
Machine Learning Scientist Intern (TikTok-Content Ecology-LLM application) - 2026 Start (PhD)
San Jose, CA · On-site
$60/hr
... learning, and recommendation algorithms. We develop cutting-edge AI capabilities that power ... Our work includes: - Short Video Content Understanding - Building multimodal AI models to analyze ...
Machine Learning Scientist Intern (TikTok-Content Ecology-LLM application) - 2026 Start (PhD)
San Jose, CA · On-site
$60/hr
... learning, and recommendation algorithms. We develop cutting-edge AI capabilities that power ... Our work includes: - Short Video Content Understanding - Building multimodal AI models to analyze ...
Machine Learning: Multimodal Foundation Models
San Francisco, CA · On-site
$200K - $350K/yr
Machine Learning: Multimodal Foundation Models We are building unified foundation models that natively reason across text, image, video, and kinematics to drive intelligent robotic policies. You will ...
Machine Learning: Multimodal Foundation Models
San Francisco, CA · On-site
$200K - $350K/yr
Machine Learning: Multimodal Foundation Models We are building unified foundation models that natively reason across text, image, video, and kinematics to drive intelligent robotic policies. You will ...
Senior Machine Learning Engineer, Data Mining
San Francisco, CA · On-site +1
$144K - $190K/yr
Omnitag, our ML-powered multimodal data mining framework, is the engine that powers this discovery. As a Senior Machine Learning Engineer on the Data Mining team, your mission is to build the "Brain ...
Quick apply
Senior Machine Learning Engineer, Data Mining
San Francisco, CA · On-site +1
$144K - $190K/yr
Omnitag, our ML-powered multimodal data mining framework, is the engine that powers this discovery. As a Senior Machine Learning Engineer on the Data Mining team, your mission is to build the "Brain ...
Multimodal Learning information
What is multimodal learning?
What is the difference between Multimodal Learning vs Data Scientist?
| Aspect | Multimodal Learning | Data Scientist |
|---|---|---|
| Required Credentials | Advanced degrees in AI, Machine Learning, or Computer Science | Bachelor's or Master's in Data Science, Statistics, or related fields |
| Work Environment | Research labs, AI development teams, academia | Business, tech companies, analytics teams |
| Industry Usage | AI research, multimedia applications, robotics | Data analysis, predictive modeling, business insights |
Multimodal Learning focuses on developing AI models that process and integrate multiple data types like images, text, and audio. Data Scientists analyze data to extract insights, build models, and support decision-making. While both roles involve data and algorithms, Multimodal Learning is specialized in AI model development for complex data integration, whereas Data Scientists work broadly across data analysis and interpretation.
What are the key skills and qualifications needed to thrive as a Multimodal Learning Specialist, and why are they important?
What are some common challenges faced by professionals working in multimodal learning roles, and how can they be addressed?
- Intern Horizon Robotics
- Internship Bioinformatics Data Scientist
- Trainee Large Language Model Llm
- Internship Machine Learning Engineer New Grad
- Freelance Computational Chemistry Phd
- Edge Ai Machine Learning
- Phd Biotech
- Temporary Retrieval Augmented Generation
- Contract Machine Learning Research Scientist
- Remote Histology Supervisor

Other
Posted 29 days ago
Job description
Focus
Multimodal Foundation Models Representation Learning Method Innovation
We are looking for strong technical builders and researchers who deeply understand foundation models and representation learning beyond simply applying existing frameworks.
Ideal candidates should have:
- Strong experimental rigor
- Solid systems and modeling intuition
- Hands-on engineering ability
- Interest in scalable multimodal AI systems for real-world autonomy
We value people who can bridge research and production, and who care about robustness, scalability, efficiency, and practical deployment in large-scale autonomous driving systems.
Responsibilities
1. Large-Scale Foundation Model Pretraining
- Develop scalable pretraining pipelines for large-scale multimodal driving data
- Design and optimize training strategies for:
- Vision-language-action models
- Video foundation models
- Long-context temporal modeling
- Multimodal representation alignment
- Improve:
- Training stability
- Data efficiency
- Scaling efficiency
- Representation robustness
- Work on distributed training systems and large-scale model optimization using frameworks such as:
- PyTorch Distributed
- DeepSpeed
- Megatron-LM
2. Representation Learning & Method Innovation
- Design and improve self-supervised and multimodal learning methods for real-world autonomous driving systems
- Conduct architecture-level research on:
- Vision Transformers (ViT)
- Video / temporal architectures
- Multimodal fusion and alignment
- Embedding and retrieval systems
- Long-context and memory-efficient architectures
- Explore and improve:
- Pretraining objectives
- Loss functions
- Training paradigms
- Generalization and robustness
- Analyze model behavior through:
- Rigorous ablation studies
- Failure case analysis
- Representation probing and evaluation
3. Efficient Foundation Models & Scalable Deployment
- Improve the efficiency, scalability, and deployability of large multimodal foundation models for real-world autonomous driving systems
- Work on areas such as:
- Model quantization
- Knowledge distillation
- Efficient attention mechanisms
- Sparse architectures and Mixture-of-Experts (MoE)
- Long-context and memory-efficient modeling
- Inference acceleration and serving optimization
- Training and inference system efficiency
- Optimize model throughput, latency, memory usage, and deployment performance for large-scale production environments
Requirements
- MS or PhD in:
- Computer Vision
- Machine Learning
- Robotics
- Computer Science
- Related fields
- Strong understanding of:
- Foundation models
- Self-supervised learning
- Representation learning
- Multimodal learning
- Large-scale pretraining
- Hands-on experience with methods such as:
- CLIP
- DINO / DINOv2
- MAE
- Contrastive learning
- Masked modeling
- MoE or scalable transformer architectures
- Experience with one or more of the following is highly valued:
- Video foundation models
- Long-context modeling
- Retrieval systems
- Efficient inference
- Distributed training
- Model compression and deployment optimization
- Strong publication record in top-tier venues is preferred:
- CVPR
- ICCV
- ECCV
- NeurIPS
- ICLR
- ICML