1

Multimodal Jobs (NOW HIRING)

Description We are looking for a Multimodal AI Researcher with a strong background in developing foundation models for generative AI and multimodal systems that integrate various types of real-time ...

Multimodal LLM Researcher

Palo Alto, CA · On-site

$300K - $400K/yr

Multimodal LLM Researcher $300,000 - $400,000 Remote, Palo Alto Full-time / Permanent DeepRec has partnered with a high-growth generative AI company (Series B, $130M+ raised). They're building ...

The Chat and Multimodal Safety team is responsible for ensuring that OpenAI's increasingly multimodal models and products behave safely across these experiences. We develop the research, training ...

Description We are looking for a Multimodal AI Researcher with a strong background in developing foundation models for generative AI and multimodal systems that integrate various types of real-time ...

Dolby is looking for a talented Senior Multimodal AI Researcher, Audio to join Dolby's research efforts and drive innovation in multimodal AI for audio applications, multimodal representations, and ...

Multimodal Technologist * Discipline: Allied Health Professional * Duration: Ongoing * Employment Type: Staff UP TO $15,000 Bonus!! Position Summary Performs a variety of radiological procedures ...

Medlivo is seeking a travel Multimodal Technologist for a travel job in Rifle, Colorado. & Requirements * Specialty: Multimodal Technologist * Discipline: Allied Health Professional * Duration: 13 ...

next page

Showing results 1-20

Multimodal information

See salary details

$55K

$87.8K

$118K

How much do multimodal jobs pay per year?

As of Aug 28, 2026, the average yearly pay for multimodal in the United States is $87,833.00, according to ZipRecruiter salary data. Most workers in this role earn between $64,000.00 and $102,500.00 per year, depending on experience, location, and employer.

What are multimodal jobs?

Multimodal jobs refer to roles that involve the integration or management of multiple modes of communication, data, or transportation. In technology, multimodal jobs often relate to developing or working with systems that process and combine different input types, such as text, images, audio, and video, to enhance user experience or improve decision-making. In logistics, multimodal jobs may involve coordinating the movement of goods using various transport methods like rail, road, sea, and air. These positions require strong organizational and communication skills, as well as familiarity with the relevant technologies or supply chain processes.

What are some common challenges faced by professionals in multimodal roles, and how can they be addressed?

Professionals working in multimodal roles often encounter the challenge of integrating data from diverse sources, such as text, images, audio, and video, to develop unified solutions. Coordinating between multidisciplinary teams—such as data scientists, engineers, and domain experts—can also be complex due to varying priorities and communication styles. To address these challenges, it is helpful to establish clear project goals, maintain open channels of communication, and leverage standardized frameworks or tools for data integration. Continuous learning and cross-functional collaboration are also essential for staying current with evolving technologies and methodologies in this rapidly growing field.

What are the key skills and qualifications needed to thrive as a Multimodal Transportation Planner, and why are they important?

To thrive as a Multimodal Transportation Planner, you need expertise in transportation planning, data analysis, and urban design, often backed by a degree in urban planning, civil engineering, or a related field. Familiarity with GIS software, transportation modeling tools, and relevant regulatory frameworks is essential. Strong communication, problem-solving, and stakeholder engagement skills help build consensus and manage complex projects. These abilities ensure effective, sustainable planning that integrates various transportation modes to meet community and environmental needs.

What is the difference between Multimodal vs Transportation Coordinator?

AspectMultimodalTransportation Coordinator
CredentialsRelevant logistics certifications, knowledge of multiple transport modesLogistics or transportation certifications, industry experience
Work EnvironmentLogistics companies, freight forwarding, supply chain managementShipping companies, freight carriers, distribution centers
Industry UsageUsed in supply chain planning involving multiple transport modesFocuses on coordinating specific shipments within transportation networks

Multimodal professionals manage logistics involving various transportation modes like rail, sea, and road, ensuring seamless integration. Transportation Coordinators focus on organizing and tracking shipments within a specific mode or carrier. While both roles require logistics knowledge and coordination skills, Multimodal roles emphasize multi-mode planning, whereas Transportation Coordinators specialize in specific transportation segments.

More about Multimodal jobs

What cities are hiring for Multimodal jobs?

Cities with the most Multimodal job openings:

What states have the most Multimodal jobs?

States with the most job openings for Multimodal jobs include:

Infographic showing various Multimodal job openings in the United States as of August 2026, with employment types broken down into 1% Internship, 1% As Needed, 91% Full Time, 3% Part Time, 2% Contract, and 2% Nights. Highlights an 83% Physical, 5% Hybrid, and 12% Remote job distribution, with an average salary of $87,833 per year, or $42.2 per hour.

Multimodal AI Researcher

Apple

Sunnyvale, CA • On-site

Full-time

Posted 17 days ago


Apple rating

8.1

Company rating: 8.1 out of 10

Based on 678 frontline employees who took The Breakroom Quiz

7th of 30 rated technology retailers


Job description

The Video Computer Vision organization is working on breakthrough technologies for future Apple products. Our team delivers cutting-edge AI, machine learning, computer vision and graphics algorithms that power technologies including human understanding, perception, digital humans, multimodal generative AI, and agents. Our algorithms ship across a range of Apple products, including iPhone and Apple Vision Pro, where our work has contributed to technologies like Personalized Spatial Audio, EyeSight, and Persona as well as future Apple products. We are an applied research group, we push the state of the art and then bring it to product. In this role, you will collaborate with world-class experts in AI, ML, Software, and Hardware to tackle fundamental challenges in human-centric solutions that will impact millions of users across Apple's ecosystem.
Description
We are looking for a Multimodal AI Researcher with a strong background in developing foundation models for generative AI and multimodal systems that integrate various types of real-time sensor data such as video and audio with other modalities like text. Our ongoing investigations include interactive models, audio-to-audio modeling and systems. You will work on hard, open research problems in multimodal generative AI and agents, and you will see that work through to real features used by millions of people. You will collaborate with others to drive data requirements, validation strategies, and key performance indicators, and conduct algorithm research and development that serves product needs.
We hire researchers who are highly motivated and deeply care about shipping. A successful candidate will stay up-to-date with the latest advancements in multimodal foundations models and applying this knowledge to drive innovation, but also take a practical approach to problem solving and software engineering.
Minimum Qualifications
BS and a minimum of 3 years relevant industry experience.
Experience building models for multimodal perception systems.
Experience working with LLMs and VLMs.
Software engineering skills and proficiency in Python and PyTorch.
Curiosity and willingness to learn new things in order to improve the quality of their solutions.
Preferred Qualifications
MS or PhD in computer vision, computer graphics, machine learning, computer science, computer engineering or related fields.
Experience in developing, training/tuning foundation models and multimodal LLMs.
Experience with training and troubleshooting generative architectures such as diffusion, reinforcement learning, flow matching or normalizing flow at scale.
Experience with real-time or streaming multimodal models.
Experience with speech understanding and generation.
Experience applying reinforcement learning to help post-train foundation models.
Excellent communication and experience working with multi-functional teams.
Self-motivated with proven track record to optimally prioritize and deliver tasks on schedule.

What Apple employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Apple logo

About Apple

Sourced by ZipRecruiter

Imagine what you could do here! At Apple, new ideas have a way of becoming extraordinary products, services, and customer experiences very quickly. Bring passion and dedication to your job and there's no telling what you could accomplish. Dynamic, intelligent people and inspiring, innovative technologies are the norm here. The people who work here have reinvented entire industries with all Apple Hardware products. The same real passion for innovation that goes into our products also applies to our practices strengthening our dedication to leave the world better than we found it.

Industry

Computer and electronic product manufacturing

Company size

10,000+ Employees

Headquarters location

Cupertino, CA, US

Year founded

1976