1

Deep Learning Quantization Jobs in Tennessee (NOW HIRING)

AI Engineer

Nashville, TN · On-site

$50K - $112K/yr

... networks and deep learning methods for advanced AI applications - Managing data quality and ... using quantization, inference acceleration, and model-routing techniques - Designing agent ...

AI Lead Engineer

Nashville, TN

$99K - $130K/yr

With deep expertise in Managed Infrastructure Services, Application Modernization, and Industry ... Promote a culture of innovation, experimentation, and learning , while maintaining strong ...

AI Lead Engineer

Nashville, TN

$99K - $130K/yr

With deep expertise in Managed Infrastructure Services, Application Modernization, and Industry ... Promote a culture of innovation, experimentation, and learning , while maintaining strong ...

AI Lead Engineer

Nashville, TN · On-site

$99K - $130K/yr

With deep expertise in Managed Infrastructure Services, Application Modernization, and Industry ... Promote a culture of innovation, experimentation, and learning , while maintaining strong ...

AI Lead Engineer

Nashville, TN · On-site

$99K - $130K/yr

... deep technical expertise in GenAI, agentic systems, and platform‑first design. • Promote a ... learning, while maintaining strong governance, operational excellence, and long‑term platform ...

Deep Learning Quantization information

What is deep learning quantization?

Deep learning quantization is the process of reducing the precision of the numbers used to represent a neural network's parameters, activations, or both. By converting the typically used 32-bit floating-point values to lower bit-width formats such as 16-bit or 8-bit integers, quantization significantly reduces the memory footprint and computational requirements of deep learning models. This technique helps deploy models efficiently on edge devices and mobile hardware while maintaining acceptable accuracy levels. Quantization is widely used in model optimization for faster inference and lower power consumption.

What are some common challenges faced when implementing deep learning quantization in production environments?

One of the main challenges in implementing deep learning quantization is balancing model accuracy with computational efficiency, as quantization can sometimes lead to a drop in model performance. Additionally, ensuring hardware compatibility and optimizing for different devices (such as CPUs, GPUs, or edge devices) can require extensive testing and tuning. Collaboration with data scientists, software engineers, and hardware specialists is often essential to successfully deploy quantized models at scale. Staying updated with the latest quantization techniques and frameworks is also important for overcoming these challenges.

What are the key skills and qualifications needed to thrive as a deep learning quantization engineer, and why are they important?

To excel as a Deep Learning Quantization Engineer, you need a strong background in machine learning, applied mathematics, and computer science, usually supported by an advanced degree in a related field. Familiarity with deep learning frameworks (such as TensorFlow or PyTorch), quantization toolkits, and hardware acceleration platforms is crucial. Analytical thinking, problem-solving, and clear technical communication are standout soft skills in this role. These abilities are essential for efficiently optimizing models for deployment on resource-constrained hardware while maintaining accuracy and performance.

What is the difference between Deep Learning Quantization vs Machine Learning Engineer?

AspectDeep Learning QuantizationMachine Learning Engineer
Required CredentialsAdvanced degrees in AI, Computer Science, or related fields; knowledge of neural networksBachelor's or Master's in CS, Data Science, or related fields; programming skills
Work EnvironmentResearch labs, AI development teams, hardware optimization settingsSoftware development teams, data-driven projects, product-focused environments
Industry UsageAI hardware optimization, model deployment, edge computingModel development, data analysis, software solutions across industries

Deep Learning Quantization focuses on reducing model size and improving inference speed through techniques like weight and activation quantization, often in hardware or embedded systems. Machine Learning Engineers develop, implement, and optimize machine learning models for various applications. While both roles require knowledge of AI and programming, Deep Learning Quantization is more specialized in model optimization techniques, whereas Machine Learning Engineers work broadly on model development and deployment.

Machine Learning Engineer at Gravity IT Resources Nashville, TN

Nashville, TN • On-site

$110 - $150/hr

Other

Posted 22 days ago


Job description

Job Description

Machine Learning Engineer

Employment Type: Full-Time

Location: Nashville, TN (hybrid)

About the Role

We’re hiring a Maching Learning Engineer to design and deploy AI systems end-to-end — from data preparation and evaluation to model fine-tuning, inference, and agentic workflows. You’ll work closely with product and engineering teams to deliver reliable, cost-effective, and scalable LLM-powered solutions on AWS.

What You’ll Do
  • End-to-End GenAI Solutions: Scope problems, choose the right approach (prompt engineering, fine-tuning, agents), implement, evaluate, and deploy.
  • Data & SQL: Write efficient SQL for analytics and data prep; manage schemas and pipelines for model training and inference.
  • Model Training & Fine-Tuning: Run supervised fine-tuning (PEFT/LoRA/QLoRA), optimize prompts, and manage experiment tracking/evaluation.
  • Agentic Systems: Build agent workflows with tool use, memory, and safety/guardrails.
  • Inference & Deployment: Package services with Docker, optimize latency and cost (batching, caching, quantization), and deploy on AWS (ECS, EKS, SageMaker, Lambda with GPU acceleration).
  • MLOps & Observability: Set up CI/CD for models/prompts; maintain offline/online evaluation pipelines, monitoring, and rollback strategies.
  • Security & Compliance: Implement data governance, PHI/PII protections, and guardrails against prompt injection and unsafe outputs.
  • Cross-Functional Collaboration: Work with product managers and engineers to align GenAI capabilities with product goals; clearly document and communicate trade-offs.
  • Production Readiness: Lead conversations around scaling, monitoring, and maintaining GenAI systems in production environments.
Minimum Qualifications
  • 5+ years of Software/ML engineering experience, including 2+ years building and deploying GenAI/LLM systems.
  • MS/PhD in Computer Science, Data Science, or equivalent experience.
  • Strong SQL and Python skills with solid software engineering fundamentals.
  • Experience with agent frameworks (LangGraph, AutoGen, CrewAI) and tool-driven agents.
  • Hands‑on with deep learning (PyTorch or TensorFlow) and LLM fine‑tuning (SFT/PEFT like LoRA/QLoRA).
  • Production experience with Docker and AWS (ECS, EKS, SageMaker, Lambda, or GPU services).
  • Experience building scalable data and model pipelines for training and deployment.
  • Familiarity with prompt engineering, evaluation frameworks (LLM‑as‑judge, metrics), and offline test harnesses.
  • Understanding of security & compliance for sensitive data (e.g., PHI/PII).
  • Excellent problem‑solving, communication, and documentation skills.
Preferred Qualifications
  • Experience with inference optimization: quantization (bitsandbytes, GPTQ/AWQ), batching, caching, or vLLM.
  • Background in healthcare, including HIPAA compliance or medical data handling.
  • Experience with experiment tracking (MLflow, W&B), CI/CD for ML, and monitoring tools (Prometheus, Grafana).
  • Familiarity with major LLM APIs and open‑source models (OpenAI, Anthropic, Llama, Mistral).
Tech Stack
  • Languages: Python, SQL
  • DL/LLM: PyTorch, TensorFlow, Hugging Face, PEFT/TRL, vLLM
  • Data: Snowflake, Postgres
  • Cloud: AWS (ECS, EKS, SageMaker, Lambda)
  • MLOps: Docker, CI/CD, MLflow, or W&B
#J-18808-Ljbffr