1

Multimodal Jobs in California (NOW HIRING)

Core Product & Platform | API - San Francisco API Multimodal builds the developer-facing products and infrastructure that bring OpenAI's image, audio, and real-time model capabilities into the world.

Developing, implementing, and enhancing multimodal foundation models. This encompasses training or fine-tuning MM-LLMs from scratch or leveraging existing technologies to optimize performance and ...

About the Team API Multimodal builds the developer-facing products and infrastructure that bring OpenAI's image, audio, and real-time model capabilities into the world. We are responsible for high ...

Developing, implementing, and enhancing multimodal foundation models. This encompasses training or fine-tuning MM-LLMs from scratch or leveraging existing technologies to optimize performance and ...

Showing results 41-60

Multimodal information

What are multimodal jobs?

Multimodal jobs refer to roles that involve the integration or management of multiple modes of communication, data, or transportation. In technology, multimodal jobs often relate to developing or working with systems that process and combine different input types, such as text, images, audio, and video, to enhance user experience or improve decision-making. In logistics, multimodal jobs may involve coordinating the movement of goods using various transport methods like rail, road, sea, and air. These positions require strong organizational and communication skills, as well as familiarity with the relevant technologies or supply chain processes.

What are some common challenges faced by professionals in multimodal roles, and how can they be addressed?

Professionals working in multimodal roles often encounter the challenge of integrating data from diverse sources, such as text, images, audio, and video, to develop unified solutions. Coordinating between multidisciplinary teams—such as data scientists, engineers, and domain experts—can also be complex due to varying priorities and communication styles. To address these challenges, it is helpful to establish clear project goals, maintain open channels of communication, and leverage standardized frameworks or tools for data integration. Continuous learning and cross-functional collaboration are also essential for staying current with evolving technologies and methodologies in this rapidly growing field.

What are the key skills and qualifications needed to thrive as a Multimodal Transportation Planner, and why are they important?

To thrive as a Multimodal Transportation Planner, you need expertise in transportation planning, data analysis, and urban design, often backed by a degree in urban planning, civil engineering, or a related field. Familiarity with GIS software, transportation modeling tools, and relevant regulatory frameworks is essential. Strong communication, problem-solving, and stakeholder engagement skills help build consensus and manage complex projects. These abilities ensure effective, sustainable planning that integrates various transportation modes to meet community and environmental needs.

What is the difference between Multimodal vs Transportation Coordinator?

AspectMultimodalTransportation Coordinator
CredentialsRelevant logistics certifications, knowledge of multiple transport modesLogistics or transportation certifications, industry experience
Work EnvironmentLogistics companies, freight forwarding, supply chain managementShipping companies, freight carriers, distribution centers
Industry UsageUsed in supply chain planning involving multiple transport modesFocuses on coordinating specific shipments within transportation networks

Multimodal professionals manage logistics involving various transportation modes like rail, sea, and road, ensuring seamless integration. Transportation Coordinators focus on organizing and tracking shipments within a specific mode or carrier. While both roles require logistics knowledge and coordination skills, Multimodal roles emphasize multi-mode planning, whereas Transportation Coordinators specialize in specific transportation segments.

What cities in California are hiring for Multimodal jobs?

Cities in California with the most Multimodal job openings:

Infographic showing various Multimodal job openings in California as of August 2026, with employment types broken down into 17% Internship, 66% Full Time, and 17% Contract. Highlights an 100% In-person job distribution.

Member of Technical Staff, Multimodal Speech

Hark

San Jose, CA • On-site

$180K - $450K/yr

Other

Re-posted 6 days ago


Job description

Member Of Technical Staff, Multimodal Speech

San Jose

Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.

We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines. While today's AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes next: agentic systems that interact naturally with people and the real world.

To get there, we're developing multimodal models and next-generation AI hardware together - designed from the ground up as a single, unified interface for a new era of intelligent systems.

About The Role

The Omni team at Hark is building the next generation of AI experiences beyond text, enabling models to understand and generate content across multiple modalities, including text, audio. Our goal is to create seamless, real-time multimodal intelligence that powers intuitive and immersive user experiences.

As part of the Omni team, you will drive the development of advanced speech and audio capabilities within multimodal foundation models. You will work across the full stack—from data and modeling to training, evaluation, and real-time serving—pushing the boundaries of speech intelligence and human-computer interaction.

Responsibilities
  • Drive research and development to advance speech and audio capabilities in multimodal models, including speech recognition, synthesis, and understanding.
  • Develop and improve large-scale speech and audio data pipelines, including data collection, filtering, alignment, and synthetic data generation.
  • Design and implement state-of-the-art models for speech and audio, including end-to-end multimodal architectures and real-time systems.
  • Build evaluation frameworks and internal benchmarks to measure speech quality, latency, robustness, and overall user experience.
  • Optimize models and systems for real-time performance, scalability, and production deployment.
  • Collaborate closely with product and engineering teams to translate research innovations into impactful, user-facing AI experiences.
Requirements
  • Proven track record of advancing speech or audio models through innovations in data, modeling, or training.
  • Strong experience in speech/audio domains such as ASR, TTS, speech-to-speech, or audio foundation models.
  • Experience with large-scale machine learning systems and distributed training.
  • Strong background in data-driven experimentation, systematic evaluation, and model iteration.
  • Strong ownership mindset and ability to drive end-to-end impact from research to production.
Bonus Qualifications
  • Familiarity with signal processing, acoustics, or audio representation learning is a plus
  • Experience with multimodal systems (speech + text, speech + vision) or real-time AI systems is a strong plus.
Compensation

The US base salary range for this full-time position is between $180,000 - $450,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.