1

Deep Learning Quantization Jobs in Charlotte, NC

... boosting, deep learning, graph-based models). Technology Familiarity * Supervised learning ... Model serving and inference optimization (vLLM-class serving, quantization). Nice-to-Have:

New

AI Engineer

Charlotte, NC · On-site

$50K - $112K/yr

... networks and deep learning methods for advanced AI applications - Managing data quality and ... using quantization, inference acceleration, and model-routing techniques - Designing agent ...

Deep Learning Quantization information

See Charlotte, NC salary details

$10.7K

$81.9K

$136.7K

How much do deep learning quantization jobs pay per year?

As of Aug 6, 2026, the average yearly pay for deep learning quantization in Charlotte, NC is $81,932.00, according to ZipRecruiter salary data. Most workers in this role earn between $70,300.00 and $135,800.00 per year, depending on experience, location, and employer.

What are the key skills and qualifications needed to thrive as a deep learning quantization engineer, and why are they important?

To excel as a Deep Learning Quantization Engineer, you need a strong background in machine learning, applied mathematics, and computer science, usually supported by an advanced degree in a related field. Familiarity with deep learning frameworks (such as TensorFlow or PyTorch), quantization toolkits, and hardware acceleration platforms is crucial. Analytical thinking, problem-solving, and clear technical communication are standout soft skills in this role. These abilities are essential for efficiently optimizing models for deployment on resource-constrained hardware while maintaining accuracy and performance.

What is the difference between Deep Learning Quantization vs Machine Learning Engineer?

AspectDeep Learning QuantizationMachine Learning Engineer
Required CredentialsAdvanced degrees in AI, Computer Science, or related fields; knowledge of neural networksBachelor's or Master's in CS, Data Science, or related fields; programming skills
Work EnvironmentResearch labs, AI development teams, hardware optimization settingsSoftware development teams, data-driven projects, product-focused environments
Industry UsageAI hardware optimization, model deployment, edge computingModel development, data analysis, software solutions across industries

Deep Learning Quantization focuses on reducing model size and improving inference speed through techniques like weight and activation quantization, often in hardware or embedded systems. Machine Learning Engineers develop, implement, and optimize machine learning models for various applications. While both roles require knowledge of AI and programming, Deep Learning Quantization is more specialized in model optimization techniques, whereas Machine Learning Engineers work broadly on model development and deployment.

What is deep learning quantization?

Deep learning quantization is the process of reducing the precision of the numbers used to represent a neural network's parameters, activations, or both. By converting the typically used 32-bit floating-point values to lower bit-width formats such as 16-bit or 8-bit integers, quantization significantly reduces the memory footprint and computational requirements of deep learning models. This technique helps deploy models efficiently on edge devices and mobile hardware while maintaining acceptable accuracy levels. Quantization is widely used in model optimization for faster inference and lower power consumption.

What are some common challenges faced when implementing deep learning quantization in production environments?

One of the main challenges in implementing deep learning quantization is balancing model accuracy with computational efficiency, as quantization can sometimes lead to a drop in model performance. Additionally, ensuring hardware compatibility and optimizing for different devices (such as CPUs, GPUs, or edge devices) can require extensive testing and tuning. Collaboration with data scientists, software engineers, and hardware specialists is often essential to successfully deploy quantized models at scale. Staying updated with the latest quantization techniques and frameworks is also important for overcoming these challenges.
What cities near Charlotte, NC are hiring for Deep Learning Quantization jobs? Cities near Charlotte, NC with the most Deep Learning Quantization job openings:

GenAI Platform / LLM Inference Optimization Engineer (Cloud)

Infosys

Charlotte, NC • On-site

Full-time

Re-posted 22 days ago


Infosys rating

7.0

Company rating: 7.0 out of 10

Based on 62 frontline employees who took The Breakroom Quiz

149th of 221 rated it services


Job description

Job Summary:
Infosys is a global leader in next-generation digital services and consulting, specializing in transforming data into actionable insights. As a Data Science Consultant 2, you will develop and optimize predictive models and algorithms, ensuring data readiness for advanced modeling and contributing to innovative analytics solutions.
Responsibilities:
• Develop data preparation tasks, while identifying patterns or anomalies.
• Ensure data readiness for advanced modeling.
• Develop models for complex use cases (e.g., forecasting models, LLM-based solutions), while refining algorithms to meet business needs, and ensure smooth deployment into scalable, production-ready solutions.
• Conduct testing and optimize algorithms for performance, reliability, and scalability, while providing guidance to team members in best practices.
• Design and develop predictive models and data-driven analyses to address business challenges.
• Build, evaluate, and deploy models, standardize code, and contribute to knowledge management.
• Leverage tools like SAS and R/Python to create reusable customizations for non-ML, ML, and deep learning algorithms, while enhancing analytics including LLMs, and create innovative, cost-effective solutions.
• Define analytics problems for projects; execute visualization, analysis, and predictive modeling under guidance.
• Proactively maintain models and implement improvements for accuracy and reliability.
• Apply governance controls to mitigate risks and ensure compliance.
• Analyze performance trends, recommend improvements, and document discrepancies for escalation.
• Maintain comprehensive documentation standards, while participating in knowledge transfer sessions.
• Participate in discussions with stakeholders to refine requirements, provide insights, and guide implementation of models.
• Apply the predefined quality measurement framework at an individual task level in the project.
• Deploy complex analytics tools or multi-system integration, while validating deployment success.
• Participate in developing scripts or templates for repeated deployments tasks.
• Contribute to analytic solutions, IP asset creation, and training initiatives.
• Contribute to thought leadership such as papers, innovative non-ML, ML, deep learning or LLM models, and proofs of concepts.
• Participate in and deliver analytics training, while contributing to content creation.
• Provide input for segment and unit-level business plans.
Qualifications:
Required:
• vLLM, TensorRT-LLM, Triton, SGLang
• Quantization (FP8/AWQ/GPTQ), tensor parallelism
• Performance benchmarking & tuning
• Kubernetes, GKE, KServe / ML serving patterns
• Helm, Operators
• GPU orchestration concepts and scheduling patterns
• GCP and/or Azure (strong hands-on)
• Terraform
• Cloud networking, landing zones, governance/org policies
• HashiCorp Vault (secrets management)
• Observability & SRE
• Prometheus/Grafana, logging, tracing
• SRE/SLO mindset, reliability engineering
• Bachelor’s degree or foreign equivalent required from an accredited institution. Will also consider three years of progressive experience in the specialty in lieu of every year of education.
• This position may require relocation and/or travel to work/project location.
• Candidates authorized to work for any employer in the United States without employer-based visa sponsorship are welcome to apply. Infosys is unable to provide immigration sponsorship for this role now or in the future.
Preferred:
• Experience in Big Data technologies (e.g., BigQuery, Hadoop).
• Expertise in ML model development, data engineering, and software engineering principles.
• Knowledge of MLOps and AI/ML deployment (e.g., SageMaker, Snowflake). Familiarity with CI/CD, DevOps, and automation tools in AI/ML contexts.
• Design and implement LLM inference serving stacks using: vLLM, TensorRT-LLM, Triton Inference Server, SGLang
• Inference optimization techniques: continuous batching, speculative decoding, KV/prefix caching
• Quantization: FP8 / AWQ / GPTQ and tuning for GPU utilization
• Build Kubernetes-based serving platforms: KServe, Kubernetes ML Serving, GKE, OpenShift (OCP) (where applicable)
• Enable GenAI platforms and RAG use cases: Integrate LLM services with RAG pipelines, Provide reusable internal libraries, templates, and developer enablement assets
• Collaborate with cross-functional teams and client stakeholders to productionize LLM workloads at scale
Company:
Infosys is a technology company that offers consulting, outsourcing, cloud infrastructure, program management, and software services. Founded in 1981, the company is headquartered in Bangalore, IND, with a team of 10001+ employees. The company is currently Late Stage.

What Infosys employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom