Job Summary:
AAMVA (American Association of Motor Vehicle Administrators) is seeking a Machine Learning (ML) Data Engineer to join their IT Division, which focuses on developing and operating information systems for motor vehicle administration. The role involves designing, building, and operationalizing ML solutions on cloud infrastructure, managing the full model lifecycle from data preparation to deployment and monitoring.
Responsibilities:
• Designing and building dataset preparation pipelines — acquiring, cleaning, transforming, and versioning data for ML training and evaluation
• Engineering features that extract meaningful signals from structured and semi-structured data sources (time-series patterns, statistical profiles, categorical encodings)
• Running structured experimentation — testing multiple algorithms against defined scenarios, measuring performance, and documenting findings
• Training, evaluating, and tuning ML models including regression, classification, clustering, anomaly detection, and ensemble methods
• Deploying models to production on cloud infrastructure and building the pipelines that keep them running (retraining, scoring, threshold management)
• Monitoring model performance in production — tracking drift, false positive rates, and detection efficacy over time
• Building and maintaining batch and streaming data pipelines using Synapse, Fabric, Spark, and Event Hubs that feed ML systems
• Writing and optimizing analytical queries (SQL, KQL, PySpark) for data exploration, statistical profiling, and real-time analysis
• Creating validation frameworks — synthetic test data generation, backtesting against historical logs, and shadow-mode evaluation
• Building dashboards and visualizations that communicate model outputs to technical and non-technical stakeholders
• Collaborating with cross-functional teams to identify ML opportunities and translate operational problems into data solutions; communicating findings, trade-offs, and model behavior clearly to technical and non-technical audiences across IT, operations, and leadership
Qualifications:
Required:
• Bachelor's degree in computer science, data science, statistics, mathematics, or related quantitative field. Equivalent work experience may be substituted
• 3–5 years of hands-on experience in data engineering, ML engineering, or applied analytics
• Hands-on cloud platform experience (Azure or AWS) building and deploying data or ML solutions on managed cloud services; specific platform less important than depth of experience
• Working knowledge of statistical foundations: distributions, variance, standard deviation, trend vs. seasonality, hypothesis testing, and how to apply them to real operational data
• Experience with the ML experiment-to-production cycle: dataset preparation, feature engineering, model training, evaluation, and deployment
• Proficiency in Python for data processing, statistical analysis, and ML model development
• Strong SQL skills with understanding of relational database fundamentals: data modeling, query optimization, indexing strategies, and how SQL Server infrastructure supports production workloads (T-SQL, stored procedures, Availability Groups)
• Experience building data pipelines that handle batch and streaming workloads
• Experience with version control systems (Git) and CI/CD practices
• Strong problem-solving skills, attention to detail, and ability to work independently on ambiguous problems
• Strong written and verbal communication skills — able to explain technical findings to non-technical stakeholders and engage productively across IT, operations, and leadership; comfort operating outside the ML silo and contributing to broader technology discussions
Preferred:
• Experience with time-series analysis, anomaly detection, or statistical process control on operational data
• Familiarity with unsupervised and semi-supervised techniques (isolation forest, clustering, ensemble methods)
• Experience building and managing ML model lifecycle on Azure (MLflow, Fabric ML, Azure ML) or AWS (SageMaker, Glue, Step Functions)
• Familiarity with KQL (Kusto Query Language) for time-series decomposition, log analytics, or real-time data exploration
• Knowledge of data modeling and dimensional modeling concepts
• Experience with synthetic test data generation and model validation frameworks
• Familiarity with operations and monitoring of mission-critical data platforms
Company:
The American Association of Motor Vehicle Administrators (AAMVA) is a tax-exempt, nonprofit organization developing model programs in motor vehicle administration, law enforcement and highway safety. Founded in 1933, the company is headquartered in Arlington, USA, with a team of 201-500 employees. The company is currently Growth Stage.