Machine Learning Engineer
Department: Machine Learning Engineer
Employment Type: Full Time
Location: Atlanta, GA
Description
We are seeking a skilled and forwardโlooking ML Engineer with experience in Large Language Models (LLMs), generative AI, and agentic architectures to join our growing R&D and Applied AI team. This role is critical in helping Oversight deliver the next generation of agentic AI systems for enterprise spend management and risk controls.
The ideal candidate has a strong foundation in machine learning, modern deep learning frameworks, and data pipelines, coupled with handsโon experience experimenting with LLMs, small language models (SLMs), multiโagent frameworks, and retrievalโaugmented generation (RAG).
You will work closely with AI/ML researchers, data engineers, and product teams to design, implement, and optimize models that power autonomous exception resolution, anomaly detection, and explainable insights. This is a handsโon engineering role where you will not only build and scale ML systems but also actively contribute to cuttingโedge applied research in agentic AI.
Key Responsibilities
- Contribute to the design, training, fineโtuning, and deployment of ML/LLM models for production.
- Implement RAG pipelines using vector databases.
- Work with frameworks like LangChain, LangGraph, MCP to prototype and optimize multiโagent workflows.
- Develop prompt engineering, optimization, and safety techniques for agentic LLM interactions.
- Integrate memory, evidence packs, and explainability modules into agentic pipelines.
- Work handsโon with multiple LLM ecosystems:
- OpenAI GPT models (GPTโ4, GPTโ4o, fineโtuned GPTs).
- Anthropic Claude (Claude 2/3 for reasoning and safetyโaligned workflows).
- Google Gemini (multimodal reasoning, advanced RAG integration).
- Meta LLaMA (fineโtuned/custom models for domainโspecific tasks).
- Collaborate with Data Engineering to build and maintain realโtime and batch data pipelines that serve ML/LLM workloads.
- Conduct feature engineering, preprocessing, and embeddings generation for structured and unstructured data.
- Implement model monitoring, drift detection, and retraining pipelines.
- Leverage cloud ML platforms (AWS SageMaker, Databricks ML) for experimentation and scaling.
- Explore and evaluate emerging LLM/SLM architectures and agent orchestration patterns.
- Experiment with generative AI and multimodal models to extend capabilities beyond text (images, structured financial data).
- Collaborate with R&D to prototype autonomous resolution agents, anomaly detection models, and reasoning engines.
- Translate research prototypes into productionโready components.
- Work crossโfunctionally with R&D, Data Science, Product, and Engineering to deliver businessโaligned AI features.
- Participate in design reviews, architecture discussions, and model evaluations.
- Document processes, experiments, and results effectively for knowledge sharing.
- Mentor junior engineers and contribute to ML engineering best practices.
Skills, Knowledge and Expertise
Required
- Bachelorโs or Masterโs degree in Computer Science, Data Science, Machine Learning, or related field.
- 3+ years of experience building and deploying ML systems.
- Proficiency in Python and libraries such as PyTorch, TensorFlow, ScikitโLearn, Hugging Face Transformers.
- Handsโon experience with LLMs/SLMs (fineโtuning, prompt design, inference optimization).
- Demonstrated experience with at least two of the following ecosystems:
- OpenAI GPT models (chat, assistants, fineโtuning).
- Anthropic Claude (safetyโfirst AI for reasoning and summarization).
- Google Gemini (multimodal reasoning, enterpriseโscale APIs).
- Meta LLaMA (openโsource, fineโtuned models).
- Familiarity with vector databases, embeddings, and RAG pipelines.
- Ability to work with structured and unstructured data at scale.
- Knowledge of SQL and distributed data frameworks (Spark, Ray).
- Strong understanding of ML lifecycle: data prep, training, evaluation, deployment, monitoring.
- Experience with agentic frameworks (LangChain, LangGraph, MCP, AutoGen).
- Knowledge of AI safety, guardrails, and explainability techniques.
- Handsโon experience deploying ML/LLM solutions in cloud environments (AWS, GCP, Azure).
- Experience with CI/CD for ML (MLOps), monitoring, and observability.
- Familiarity with anomaly detection, fraud/risk modeling, or behavioral analytics.
- Contributions to openโsource AI/ML projects or publications in applied ML research.
#J-18808-Ljbffr