1

Multimodal Learning Jobs in Washington, DC (NOW HIRING)

Software Engineer II

Herndon, VA · On-site

$100K - $137K/yr

In this role, you will develop and optimize advanced machine learning solutions supporting multimodal artificial intelligence and computer vision applications for national security missions. Working ...

Software Engineer II

Herndon, VA · On-site

$100K - $137K/yr

In this role, you will develop and optimize advanced machine learning solutions supporting multimodal artificial intelligence and computer vision applications for national security missions. Working ...

... multimodal intelligence workflows. This role is execution-focused and suited for engineers with strong foundations in computer vision, image processing, and machine learning who want to apply their ...

Lead AI Engineer

Rockville, MD · On-site

$104K - $137K/yr

The role will architect multimodal ingestion, retrieval-augmented generation, and LLM-driven ... Experience Requirements * 8+ years in software engineering or machine learning with production AI ...

Lead AI Engineer

Rockville, MD · On-site

$104K - $137K/yr

The role will architect multimodal ingestion, retrieval-augmented generation, and LLM-driven ... Experience Requirements: * 8+ years in software engineering or machine learning with production AI ...

Lead AI Engineer

Rockville, MD · On-site

$104K - $137K/yr

The role will architect multimodal ingestion, retrieval-augmented generation, and LLM-driven ... Experience Requirements: * 8+ years in software engineering or machine learning with production AI ...

... learning. * Previous experience negotiating milestone-based Statements of Work contracts. Preferred Qualifications * Demonstrated expertise in applied AI/ML, generative and multimodal AI, enterprise ...

... learning. * Previous experience negotiating milestone-based Statements of Work contracts. Preferred Qualifications * Demonstrated expertise in applied AI/ML, generative and multimodal AI, enterprise ...

Technology Architect - AI

Mclean, VA · On-site

$138K - $180K/yr

... learning. * Previous experience negotiating milestone-based Statements of Work contracts. Preferred Qualifications * Demonstrated expertise in applied AI/ML, generative and multimodal AI, enterprise ...

Associate Professor

Washington, DC · On-site

$100K - $115K/yr

Demonstrated expertise in AI, machine learning, or multimodal modeling. * A record of (or potential for) impactful research, evidenced by publications, grants, or industry partnerships. * Experience ...

Showing results 21-40

Multimodal Learning information

See Washington, DC salary details

$23.8K

$69.9K

$129.7K

How much do multimodal learning jobs pay per year?

As of Aug 10, 2026, the average yearly pay for multimodal learning in Washington, DC is $69,872.00, according to ZipRecruiter salary data. Most workers in this role earn between $46,400.00 and $81,500.00 per year, depending on experience, location, and employer.

What is multimodal learning?

Multimodal learning is an area of machine learning that involves integrating and processing information from multiple types of data, such as text, images, audio, and video. The goal is to create models that can understand and make predictions based on more than one data modality, similar to how humans use various senses. This approach is used in applications like speech recognition with visual cues, image captioning, and video analysis. By combining different data types, multimodal learning systems can achieve better accuracy and more robust understanding.

What is the difference between Multimodal Learning vs Data Scientist?

AspectMultimodal LearningData Scientist
Required CredentialsAdvanced degrees in AI, Machine Learning, or Computer ScienceBachelor's or Master's in Data Science, Statistics, or related fields
Work EnvironmentResearch labs, AI development teams, academiaBusiness, tech companies, analytics teams
Industry UsageAI research, multimedia applications, roboticsData analysis, predictive modeling, business insights

Multimodal Learning focuses on developing AI models that process and integrate multiple data types like images, text, and audio. Data Scientists analyze data to extract insights, build models, and support decision-making. While both roles involve data and algorithms, Multimodal Learning is specialized in AI model development for complex data integration, whereas Data Scientists work broadly across data analysis and interpretation.

What are the key skills and qualifications needed to thrive in multimodal learning, and why are they important?

To excel as a Multimodal Learning Specialist, you need a solid background in machine learning, data science, and computer vision, often supported by an advanced degree in a related field. Familiarity with deep learning frameworks like TensorFlow or PyTorch, experience integrating data from diverse sources (e.g., text, audio, images), and knowledge of relevant algorithms are crucial. Strong problem-solving abilities, creativity, and effective collaboration are standout soft skills for this role. These competencies are vital for developing innovative models that can process and interpret complex, multi-source data to drive impactful AI solutions.

What are some common challenges faced by professionals working in multimodal learning roles, and how can they be addressed?

Professionals in multimodal learning frequently encounter challenges related to integrating and aligning data from multiple sources, such as text, images, audio, or video. Ensuring data quality and consistency across modalities can be complex, and developing models that effectively combine heterogeneous information often requires advanced technical skills and innovative thinking. Collaboration with domain experts and other data scientists is key to overcoming these obstacles, as is staying up to date with the latest research and tools in machine learning. Regular team meetings and cross-disciplinary workshops can help foster a collaborative environment and promote knowledge sharing.
What are popular job titles related to Multimodal Learning jobs in Washington, DC? For Multimodal Learning jobs in Washington, DC, the most frequently searched job titles are:
What job categories do people searching Multimodal Learning jobs in Washington, DC look for? The top searched job categories for Multimodal Learning jobs in Washington, DC are:
Infographic showing various Multimodal Learning job openings in Washington, DC as of June 2026, with employment types broken down into 1% As Needed, 94% Full Time, 3% Part Time, 1% Temporary, and 1% Contract. Highlights an 91% Physical, 1% Hybrid, and 8% Remote job distribution, with an average salary of $69,872 per year, or $33.6 per hour.

Machine Learning Engineer (GoLang)

Blueface Ltd

Washington, DC • On-site

$142.65 - $213.98/hr

Other

Posted 5 days ago


Job description

Company Overview

Make your mark at Comcast – a Fortune 30 global media and technology company. From the connectivity and platforms we provide to the content and experiences we create, we reach hundreds of millions of customers, viewers, and guests worldwide. Join our award‑winning technology team that turns big ideas into cutting‑edge products, platforms, and solutions that our customers love.

Job Summary

Multimodal Analysis Framework (MAF) is an end‑to‑end platform designed to process diverse content sources—including video, images, audio, and documents—to generate rich, structured metadata. The platform unifies multiple ML/AI models to extract curated insights at scale, tailored to specific business needs. MAF supports both on‑demand workloads (batch uploads, ad‑hoc analysis) and real‑time streaming workflows, enabling continuous metadata generation for live content streams.

Customers can define their metadata requirements—such as entity extraction, scene segmentation, object detection, transcription, summarization, or multimodal correlation—and the framework orchestrates the appropriate models and toolchains to deliver high‑quality outputs. Through flexible APIs and UI‑based workflows, customers and internal teams can visualize metadata, trigger enrichment, monitor processing, and integrate results into downstream applications. The platform emphasizes modularity, scalability, and extensibility to support new ML models, LLM‑based agents, and cross‑modal inference as use cases evolve.

Role Overview

We are looking for a mid‑level Backend Engineer to join our Machine Learning Platform team. This role focuses on building scalable backend systems that power ML workloads, including video, image, and document processing, and enable LLM‑driven applications through agents and MCP servers. You will work primarily in Golang, deploy and operate services on Kubernetes, manage infrastructure with Terraform, and build on AWS. A core part of the role is designing platform capabilities that allow LLMs to safely and reliably interact with tools, data, and services via agent frameworks and MCP servers.

Primary Responsibilities
  • Design, build, and maintain high‑performance backend services in Golang for ML and AI platform use cases.
  • Develop REST and gRPC APIs for inference, processing pipelines, orchestration, and platform services.
  • Implement asynchronous and distributed processing patterns (workers, queues, event‑driven systems).
  • Ensure backend services meet production standards for scalability, reliability, and security.
  • Build and operate backend systems supporting video processing (frame extraction, metadata generation, embeddings, indexing).
  • Build and operate backend systems supporting image processing (OCR, classification, detection, embedding generation).
  • Build and operate backend systems supporting document processing (parsing, layout analysis, chunking, OCR, retrieval pipelines).
  • Integrate ML inference services into backend workflows with attention to latency, throughput, and cost.
  • Work closely with ML engineers and data scientists to productionize models and pipelines.
  • Build LLM‑enabled backend services using structured prompting, tool/function calling, and retrieval‑augmented generation (RAG).
  • Design and implement agentic workflows (multi‑step reasoning, tool orchestration, retries, guardrails).
  • Develop and operate MCP servers that expose internal platform capabilities (search, retrieval, processing, data access) to LLM‑based applications.
  • Enforce security, access control, and observability for agent and MCP interactions.
  • Design and maintain vector‑based retrieval systems using Milvus.
  • Implement embedding ingestion, indexing, and query pipelines at scale.
  • Optimize retrieval quality, latency, and relevance for downstream LLM applications.
  • Deploy and operate backend and ML services on Kubernetes (scaling, rollouts, resource management).
  • Use Terraform for infrastructure provisioning and continuous delivery of cloud resources.
  • Build and operate primarily on AWS, leveraging services such as compute, networking, IAM, object storage, managed Kubernetes, and logging/monitoring services.
  • Implement observability using logs, metrics, and traces; define SLOs and alerts.
  • Write automated tests (unit, integration) and contribute to CI/CD pipelines.
  • Participate in on‑call rotations and incident response; drive post‑incident improvements.
Qualifications
  • 3–6 years of professional software engineering experience.
  • Strong backend engineering experience with Golang.
  • Experience building and operating APIs (REST and/or gRPC) in production.
  • Hands‑on experience with Kubernetes in production environments.
  • Experience using Terraform for infrastructure provisioning and deployment.
  • Solid working knowledge of AWS cloud services and core architectural concepts.
  • Experience building or supporting ML processing pipelines (video, image, or document).
  • Practical experience using LLMs in production systems.
  • Experience developing agents and/or MCP servers, or equivalent tool‑integration platforms.
  • Experience with Milvus or other vector databases in production. (Preferred)
  • Familiarity with GPU‑backed workloads and ML inference optimization.
  • Experience with messaging/streaming systems (Kafka, SQS, SNS, etc.).
  • Knowledge of secure system design for AI platforms (IAM, secrets management, least‑privilege access).
  • Experience working on internal developer platforms or ML infrastructure teams.
Education

Bachelor's Degree (preferred). Comcast may consider applicants who hold some combination of coursework and experience, or who have extensive related professional experience.

Compensation

Primary Location Pay Range: $142,651.46 - $213,977.19. The application window is 30 days from the date the job is posted, unless the number of applicants requires it to close sooner or later.

Benefits

We provide best‑in‑class benefits to eligible employees. Our benefits are tailored to support you physically, financially, and emotionally through major milestones and everyday life. To learn more about our benefits, please visit the compensation and benefits summary on our careers site.

Application Statement

Comcast is an equal opportunity workplace. We will consider all qualified applicants for employment without regard to race, color, religion, age, sex, sexual orientation, gender identity, national origin, disability, veteran status, genetic information, or any other basis protected by applicable law.

#J-18808-Ljbffr