1

Multimodal Jobs in Washington (NOW HIRING)

Multi-Modal AI & Image Search You will support multimodal AI systems that combine vision models with LLMs, embeddings, and retrieval pipelines to enable natural-language search and reasoning over ...

Image & Computer Vision AI Engineer

Reston, VA · On-site

$119K - $143K/yr

Multi-Modal AI & Image Search You will support multimodal AI systems that combine vision models with LLMs, embeddings, and retrieval pipelines to enable natural-language search and reasoning over ...

Vision-language models (VLMs), Multimodal learning, Reasoning models, Large language models (LLMs), Computer vision or geospatial AI. • Strong programming skills in Python, with experience using ...

Software Engineer II

Herndon, VA · On-site

$100K - $137K/yr

In this role, you will develop and optimize advanced machine learning solutions supporting multimodal artificial intelligence and computer vision applications for national security missions. Working ...

Transportation Planning Intern - Fall 2026

Washington, DC · On-site

$17.50 - $22.75/hr

... multimodal transportation, transportation demand management, corridor planning, transit-oriented development, and related initiatives. Responsibilities may include collecting and analyzing ...

Software Engineer II

Herndon, VA · On-site

$100K - $137K/yr

In this role, you will develop and optimize advanced machine learning solutions supporting multimodal artificial intelligence and computer vision applications for national security missions. Working ...

Applied AI Scientist

Herndon, VA · On-site

$146K - $244K/yr

Productionize reasoning models, vision-language models (VLMs), and multimodal AI systems that combine imagery, geospatial signals, and structured data. * Architect enterprise-grade training and ...

Transportation Planning Intern - Fall 2026

Washington, DC · On-site

$17.50 - $22.75/hr

... multimodal transportation, transportation demand management, corridor planning, transit-oriented development, and related initiatives. Responsibilities may include collecting and analyzing ...

Applied AI Scientist

Herndon, VA · On-site

$146K - $244K/yr

Productionize reasoning models, vision-language models (VLMs), and multimodal AI systems that combine imagery, geospatial signals, and structured data. * Architect enterprise-grade training and ...

Showing results 21-40

Multimodal information

What are the key skills and qualifications needed to thrive as a Multimodal Transportation Planner, and why are they important?

To thrive as a Multimodal Transportation Planner, you need expertise in transportation planning, data analysis, and urban design, often backed by a degree in urban planning, civil engineering, or a related field. Familiarity with GIS software, transportation modeling tools, and relevant regulatory frameworks is essential. Strong communication, problem-solving, and stakeholder engagement skills help build consensus and manage complex projects. These abilities ensure effective, sustainable planning that integrates various transportation modes to meet community and environmental needs.

What are multimodal jobs?

Multimodal jobs refer to roles that involve the integration or management of multiple modes of communication, data, or transportation. In technology, multimodal jobs often relate to developing or working with systems that process and combine different input types, such as text, images, audio, and video, to enhance user experience or improve decision-making. In logistics, multimodal jobs may involve coordinating the movement of goods using various transport methods like rail, road, sea, and air. These positions require strong organizational and communication skills, as well as familiarity with the relevant technologies or supply chain processes.

What are some common challenges faced by professionals in multimodal roles, and how can they be addressed?

Professionals working in multimodal roles often encounter the challenge of integrating data from diverse sources, such as text, images, audio, and video, to develop unified solutions. Coordinating between multidisciplinary teams—such as data scientists, engineers, and domain experts—can also be complex due to varying priorities and communication styles. To address these challenges, it is helpful to establish clear project goals, maintain open channels of communication, and leverage standardized frameworks or tools for data integration. Continuous learning and cross-functional collaboration are also essential for staying current with evolving technologies and methodologies in this rapidly growing field.

What is the difference between Multimodal vs Transportation Coordinator?

AspectMultimodalTransportation Coordinator
CredentialsRelevant logistics certifications, knowledge of multiple transport modesLogistics or transportation certifications, industry experience
Work EnvironmentLogistics companies, freight forwarding, supply chain managementShipping companies, freight carriers, distribution centers
Industry UsageUsed in supply chain planning involving multiple transport modesFocuses on coordinating specific shipments within transportation networks

Multimodal professionals manage logistics involving various transportation modes like rail, sea, and road, ensuring seamless integration. Transportation Coordinators focus on organizing and tracking shipments within a specific mode or carrier. While both roles require logistics knowledge and coordination skills, Multimodal roles emphasize multi-mode planning, whereas Transportation Coordinators specialize in specific transportation segments.

Infographic showing various Multimodal job openings in Washington as of August 2026, with employment types broken down into 1% Internship, 1% As Needed, 91% Full Time, 3% Part Time, 3% Contract, and 1% Nights. Highlights an 79% Physical, 5% Hybrid, and 16% Remote job distribution.

Image & Computer Vision AI Engineer

Babel Street

Reston, VA

Other

Medical, Dental, Vision, Life, Retirement

Re-posted 23 days ago


Job description

ROLE SUMMARY:

As an Engineer on the Image & Computer Vision AI team, you will play a hands-on role in developing and deploying computer vision capabilities that support Babel Street's intelligence applications. You will build systems that extract, analyze, and reason over visual data-enabling facial matching, object and scene understanding, geolocation and location inference from imagery, and multimodal intelligence workflows. 

This role is execution-focused and suited for engineers with strong foundations in computer vision, image processing, and machine learning who want to apply their skills to real-world, mission-driven problems. You will work closely with AI, Product, and Engineering teams to deliver reliable, scalable, and cost-efficient vision capabilities, including integration with multimodal LLM systems that allow users to search and reason over images using natural language. 

This is a hybrid role to be based out of either our Reston, VA/Washington DC office or our Somerville MA office.

ROLE FOCUS;

This role spans three practical execution areas: 

Computer Vision & Image Analytics 

You will implement and operate image analytics pipelines that support facial matching, object detection, scene understanding, and image similarity. This includes image preprocessing, feature extraction, model inference, evaluation, and performance optimization to meet mission-grade accuracy and latency requirements. 

Geospatial & Location Inference from Imagery 

You will contribute to capabilities that infer location, context, or environmental attributes from imagery-leveraging visual cues, metadata, and learned representations. This includes supporting image-based geolocation, landmark recognition, and contextual scene analysis used in intelligence workflows. 

Multi-Modal AI & Image Search 

You will support multimodal AI systems that combine vision models with LLMs, embeddings, and retrieval pipelines to enable natural-language search and reasoning over images and image collections. You will help integrate visual understanding into broader intelligence applications and workflows. 

KEY RESPONSIBILITIES:

  • Build and maintain computer vision pipelines for image ingestion, preprocessing, inference, and evaluation. 
  • Implement facial matching, and identity-related vision workflows in accordance with accuracy, safety, and compliance requirements. 
  • Develop and support object detection, image similarity, and scene understanding models. 
  • Contribute to image-based geolocation and location inference capabilities using visual features and contextual signals. 
  • Support multimodal AI workflows that combine image embeddings with LLM-based search and reasoning. 
  • Write clean, maintainable Python code and contribute to production services and APIs. 
  • Assist with model evaluation, bias testing, and accuracy monitoring for vision systems. 
  • Optimize inference pipelines for performance, scalability, and cost efficiency (GPU usage, batching, model selection). 
  • Collaborate with Product and Engineering teams to integrate vision capabilities into user-facing intelligence applications. 

QUALIFICATIONS: 

Required 

  • 3+ years of experience in computer vision, image processing, or applied machine learning. 
  • Hands-on experience with computer vision models and techniques (e.g., CNNs, transformers for vision, feature embeddings). 
  • Experience building or integrating image analytics such as facial recognition, object detection, or image similarity. 
  • Strong programming skills in Python; experience with common CV/ML libraries (PyTorch, TensorFlow, OpenCV, etc.). 
  • Solid understanding of machine learning fundamentals, model evaluation, and performance tradeoffs. 
  • Experience working with large image datasets and production ML pipelines. 
  • Ability to work collaboratively in a fast-moving, mission-driven engineering environment. 

Preferred 

  • Experience with facial matching or biometric systems in regulated or high-stakes environments.
  • Experience with image-based geolocation or scene/location inference.
  • Familiarity with multimodal AI systems, including combining vision models with LLMs or natural-language search. 

EDUCATION:

Bachelor's degree in Computer Science, Engineering, Data Science, or a related technical field required. 
Advanced degree is a plus but not required. 

Benefits at Babel Street (just to name a few...)

  • Health Benefits: Babel Street covers 85-100% monthly premium costs for Medical, Dental, Vision, Life & Disability insurances - for you and your family!
  • Retirement Plans: Babel Street offers both a Traditional and Roth 401(K) with a very competitive match.
  • Unlimited Flexible Leave: We trust our employees to manage their own time and balance their personal and work lives.
  • Holidays: Babel Street provides employees with 12 paid Federal Holidays
  • Tuition Reimbursement: We are committed to investing in our employees. One way we do that is with our Tuition Reimbursement Program for continuing education.                 

Babel Street is an equal opportunity/affirmative action employer. All qualified applicants will receive consideration for employment without regard to sex, gender identity, sexual orientation, race, color, religion, national origin, disability, protected Veteran status, age, or any other characteristic protected by law. Further, Babel Street will not discriminate against applicants for inquiring about, discussing or disclosing their pay or, in certain circumstances, the pay of their coworker, Pay Transparency Nondiscrimination. In addition, Babel Street's policy is to provide reasonable accommodation to qualified employees who have protected disabilities to the extent required by applicable laws, regulations and ordinances where a particular employee works. Upon request, we will provide you with more information about such accommodations.