1

Multimodal Learning Jobs in Atlanta, GA (NOW HIRING)

... multimodal foundation models for XR, with audio focus. Develop and combine deep learning methodologies with perceptually relevant signal processing and metrics. Partner with ATG researchers, develop ...

Traffic Engineer

Duluth, GA ยท On-site

$80K - $109K/yr

Conduct multimodal analyses for vehicles, pedestrian, bicyclists and transit. * Prepare technical ... Learning & Development: We provide clear career paths, learning resources and development programs ...

Advanced Analytics & Machine Learning * Design, develop, and deploy statistical, predictive, and ... Stay current with cutting-edge advancements in foundation models, multimodal AI, and agentic ...

... multimodal machinegenerated data - including logs, time series, traces, and events! We combine deep ... Large-scale graph representation learning and Graph Neural Networks (GNNs) (e.g., GCN/GAT/GraphSAGE ...

Showing results 21-40

Multimodal Learning information

See Atlanta, GA salary details

$20.2K

$59.3K

$110.1K

How much do multimodal learning jobs pay per year?

As of Sep 2, 2026, the average yearly pay for multimodal learning in Atlanta, GA is $59,327.00, according to ZipRecruiter salary data. Most workers in this role earn between $39,400.00 and $69,200.00 per year, depending on experience, location, and employer.

What is multimodal learning?

Multimodal learning is an area of machine learning that involves integrating and processing information from multiple types of data, such as text, images, audio, and video. The goal is to create models that can understand and make predictions based on more than one data modality, similar to how humans use various senses. This approach is used in applications like speech recognition with visual cues, image captioning, and video analysis. By combining different data types, multimodal learning systems can achieve better accuracy and more robust understanding.

What are the key skills and qualifications needed to thrive in multimodal learning, and why are they important?

To excel as a Multimodal Learning Specialist, you need a solid background in machine learning, data science, and computer vision, often supported by an advanced degree in a related field. Familiarity with deep learning frameworks like TensorFlow or PyTorch, experience integrating data from diverse sources (e.g., text, audio, images), and knowledge of relevant algorithms are crucial. Strong problem-solving abilities, creativity, and effective collaboration are standout soft skills for this role. These competencies are vital for developing innovative models that can process and interpret complex, multi-source data to drive impactful AI solutions.

What are some common challenges faced by professionals working in multimodal learning roles, and how can they be addressed?

Professionals in multimodal learning frequently encounter challenges related to integrating and aligning data from multiple sources, such as text, images, audio, or video. Ensuring data quality and consistency across modalities can be complex, and developing models that effectively combine heterogeneous information often requires advanced technical skills and innovative thinking. Collaboration with domain experts and other data scientists is key to overcoming these obstacles, as is staying up to date with the latest research and tools in machine learning. Regular team meetings and cross-disciplinary workshops can help foster a collaborative environment and promote knowledge sharing.

What is the difference between Multimodal Learning vs Data Scientist?

AspectMultimodal LearningData Scientist
Required CredentialsAdvanced degrees in AI, Machine Learning, or Computer ScienceBachelor's or Master's in Data Science, Statistics, or related fields
Work EnvironmentResearch labs, AI development teams, academiaBusiness, tech companies, analytics teams
Industry UsageAI research, multimedia applications, roboticsData analysis, predictive modeling, business insights

Multimodal Learning focuses on developing AI models that process and integrate multiple data types like images, text, and audio. Data Scientists analyze data to extract insights, build models, and support decision-making. While both roles involve data and algorithms, Multimodal Learning is specialized in AI model development for complex data integration, whereas Data Scientists work broadly across data analysis and interpretation.

What cities near Atlanta, GA are hiring for Multimodal Learning jobs?

Cities near Atlanta, GA with the most Multimodal Learning job openings:

Research Scientist- Spatial Audio & AI

Dolby Laboratories, Inc.

Atlanta, GA โ€ข On-site

Full-time

Re-posted 24 days ago


Job description

Join the leader in entertainment innovation and help us design the future. At Dolby, science meets art, and high tech means more than computer code. As a member of the Dolby team, you'll see and hear the results of your work everywhere, from movie theaters to smartphones. We continue to revolutionize how people create, deliver, and enjoy entertainment worldwide. To do that, we need the absolute best talent. We're big enough to give you all the resources you need, and small enough so you can make a real difference and earn recognition for your work. We offer a collegial culture, challenging projects, and excellent compensation and benefits,ย not to mention a Flex Work approach that is truly flexible to support where, when, and how you do your best work.

The Advanced Technology Group (ATG) is the research division of the company. ATG's mission is to look ahead, deliver insights, and innovate technological solutions that will fuel Dolby's continued growth. Our researchers have a broad range of expertise related to computer science and electrical engineering, such as AI/ML, algorithms, digital signal processing, audio engineering, image processing, computer vision, data science & analytics, distributed systems, cloud, edge & mobile computing, computer networking, and IoT.

We are seeking a talented Senior Spatial Audio and Multimodal AI Researcher to join the Perceptual and Interactive Multimedia Computing team in the Multimodal Experiences Lab.

We are a key research team within Dolby's Advanced Technology Group, focused on creating cutting edge multimodal technologies that drive next generation experiences. We're looking for skilled researchers who are excited to advance the state of the art in technologies of interest to Dolby as well as the human society at large, in particular, in the area of developing AI solutions for Spatial Media/XR audio content creation workflow.

We welcome the opportunity to have you join our growing Atlanta Advanced Technology Research team.

What You Will Accomplish

  • Design AI models for spatial audio content creation and audio engineering workflow
  • Create multimodal foundation models for XR, with audio focus.
  • Develop and combine deep learning methodologies with perceptually relevant signal processing and metrics.
  • Partner with ATG researchers, develop solutions for the relevant applications.ย 

What you need to succeed

Competencies:

  • Technical depth:ย Necessary technical knowledge to create new AI algorithms and multimodal models with an audio focus. ย Solid knowledge of Audio, ML and AI fundamentals.
  • Explore new technologies: Openness to learn new skills, work with cutting-edge technologies, and innovate in new areas.
  • Invent & Innovate:ย Develop know-how, algorithms and software tools with both a short and long-term focus that further strengthen Dolby as a world leader for sight and sound experiences associated with digital content consumption. Then influence and collaborate with business group ย partners putting the technology into production.
  • Work with a sense of Urgency:ย Respond aggressively to changing trends and new technologies and creates new algorithms to capitalize on them. Take appropriate risks to be ahead of the competition and the market.
  • Collaborate:ย Collaborate with and influence peers in developing industry-leading technologies. Work with external trendsetters and technology drivers in academia and in partner enterprises.

Desired Background:

  • PhD in Computer Science, Electrical and Computer Engineering, or similar fields
  • Proven ability to pursue new areas of multimodal research for Audio, AI, and signal analysis, and demonstrate results through projects, prototypes, patent filings, and papers in peer reviewed journals and conferences
  • High comfort level in creating Algorithms in Python
  • Solid knowledge on audio signal analysis, spatial analysis, creation, and generation
  • Solid knowledge on AI/ML, e.g. large language model and generative AI
  • Familiarity with deep learning frameworks, e.g., TensorFlow, PyTorch, etc..
  • You have effective problem-solving, partnership,communication and presentation skillsย 

Learn more about our innovative research:ย https://www.dolby.com/about/innovation/empowering/

The Atlanta Area base salary range for this full-time position isย $140,000-$170,000,ย which can vary if outside this location,ย plus bonus, benefits, profit sharing and equity compensation.. Our salary ranges are determined by role, level, and location. Within the range, individual pay is determined by work location and additional factors, including job-related skills, competencies, experience, market demands, internal parity, and relevant education or training. Your recruiter can share more about the specific salary range and perks and benefits for your location during the hiring process.

Dolby will consider qualified applicants with criminal histories in a manner consistent with the requirements of San Francisco Police Code, Article 49, and Administrative Code, Article 12

Equal Employment Opportunity:
Dolby is proud to be an equal opportunity employer. Our success depends on the combined skills and talents of all our employees. We are committed to making employment decisions without regard to race, religious creed, color, age, sex, sexual orientation, gender identity, national origin, religion, marital status, family status, medical condition, disability, military service, pregnancy, childbirth and related medical conditions or any other classification protected by federal, state, and local laws and ordinances.