Title: AI/Machine Learning Engineer โ Vision Language Models / Multimodal AI (NGA)
Location: Springfield or Herndon, VA (onsite)
Clearance: TS/SCI (CI Poly preferred)
Position Type: Full-Time, Direct Hire
Pay: $175,000 to $250,000 for an SME Company: The name of our partner organization will be disclosed during the interview
process. This is not a direct role with LaunchCode; it is a position through LaunchCode,
working with one of our partner companies. Disclaimer: We are unable to provide work sponsorship for this role Overview: Weโre hiring a AI/Machine Learning Engineer with strong experience in multimodal AI and
large-scale model training to support advanced vision-language initiatives in a secure
government environment. This role will focus on fine-tuning Vision Language Models
(VLMs) on domain-specific geospatial imagery, building scalable AWS training
infrastructure, and developing evaluation frameworks for image understanding and spatial
reasoning. Ideal candidates will have deep experience with PyTorch, HuggingFace,
distributed training, and computer vision, along with the ability to optimize and deploy
multimodal models in mission-critical environments. Huge plus for candidates who have hands-on experience taking multimodal models such
as CLIP, LLaVA, Qwen-VL, or similar Vision Language Models and fine-tuning them on
classified or mission-specific imagery datasets. The ideal candidate can build the AWS
infrastructure needed to train and scale these models, evaluate performance
improvements across real-world use cases, and deploy solutions into secure government
or air-gapped environments. Key Responsibilities: โข Design and execute fine-tuning pipelines for Vision Language Models (VLMs) using domain-specific imagery datasets โข Handle data preprocessing, training orchestration, and hyperparameter optimization for multimodal models โข Build evaluation frameworks for image understanding, visual question answering, and spatial reasoning tasks โข Develop scalable AWS-based ML infrastructure using SageMaker and GPU-enabled EC2 for distributed training โข Create data pipelines for curating, annotating, and transforming geospatial imagery into model-ready datasets โข Partner with applied scientists and architects on model architecture improvements, LoRA/QLoRA strategies, and inference optimization, Required Qualifications: โข Active TS/SCI with CI Poly โข 5+ years of machine learning engineering experience focused on deep learning โข 1+ year of hands-on experience fine-tuning foundation models (LLMs or VLMs) โข Experience with LoRA, QLoRA, adapters, supervised fine-tuning, instruction tuning, and RLHF/DPO โข 4+ years of advanced Python development for ML workloads โข Strong PyTorch and HuggingFace experience (Transformers, PEFT, Datasets, Accelerate) โข Experience with distributed training frameworks such as DeepSpeed, FSDP, or Megatron โข 3+ years working with computer vision or multimodal models โข Familiarity with vision transformer architectures (ViT, CLIP, LLaVA, etc.) โข Experience processing and augmenting image datasets at scale โข 3+ years with AWS ML infrastructure including SageMaker, EC2 GPU environments, and S3 โข Experience with ML evaluation pipelines, benchmarking, metrics, and result analysis โข Strong software engineering fundamentals including version control, testing, and CI/CD Preferred Qualifications: โข 2+ years working with geospatial or remote sensing imagery โข Experience with EO or SAR satellite imagery โข Understanding of geospatial metadata, coordinate systems, and imagery preprocessing โข Experience with model quantization / inference optimization (vLLM, TensorRT, ONNX) โข MLOps tooling experience (MLflow, Weights & Biases, SageMaker Experiments) โข Familiarity with annotation tools and active learning workflows โข Containerized ML experience with Docker / ECR / ECS / EKS โข Experience supporting ATO processes and NIST 800-53 compliance โข Experience deploying in air-gapped/disconnected environments โข Familiarity with multimodal evaluation benchmarks (MMMU, MMBench, GQA) โข Publications or contributions in computer vision, multimodal AI, or VLMs โข Synthetic data generation experience for training augmentation