... ML, ML, and deep learning algorithms, while enhancing analytics including LLMs, and create ... Required : • vLLM, TensorRT-LLM, Triton, SGLang • Quantization (FP8/AWQ/GPTQ), tensor ...
... ML, ML, and deep learning algorithms, while enhancing analytics including LLMs, and create ... Required : • vLLM, TensorRT-LLM, Triton, SGLang • Quantization (FP8/AWQ/GPTQ), tensor ...
... ML, ML, and deep learning algorithms, while enhancing analytics including LLMs, and create ... Required Skill and Experience • vLLM, TensorRT-LLM, Triton, SGLang • Quantization (FP8/AWQ/GPTQ ...
... ML, ML, and deep learning algorithms, while enhancing analytics including LLMs, and create ... Required Skill and Experience • vLLM, TensorRT-LLM, Triton, SGLang • Quantization (FP8/AWQ/GPTQ ...
... deep learning models including LLMs using predefined processes and tools like SAS and R/ Python ... Proficiency in model quantization techniques (FP8, AWQ, GPTQ) and performance tuning • Experience ...
... deep learning models including LLMs using predefined processes and tools like SAS and R/ Python ... Proficiency in model quantization techniques (FP8, AWQ, GPTQ) and performance tuning • Experience ...
... ML, ML, and deep learning algorithms, while enhancing analytics including LLMs, and create ... Required Skill and Experience • vLLM, TensorRT-LLM, Triton, SGLang • Quantization (FP8/AWQ/GPTQ ...
... ML, ML, and deep learning algorithms, while enhancing analytics including LLMs, and create ... Required Skill and Experience • vLLM, TensorRT-LLM, Triton, SGLang • Quantization (FP8/AWQ/GPTQ ...
... deep learning models including LLMs using predefined processes and tools like SAS and R/ Python ... Proficiency in model quantization techniques (FP8, AWQ, GPTQ) and performance tuning Experience ...
... deep learning models including LLMs using predefined processes and tools like SAS and R/ Python ... Proficiency in model quantization techniques (FP8, AWQ, GPTQ) and performance tuning Experience ...
... deep learning models including LLMs using predefined processes and tools like SAS and R/ Python ... Proficiency in model quantization techniques (FP8, AWQ, GPTQ) and performance tuning Experience ...
... deep learning models including LLMs using predefined processes and tools like SAS and R/ Python ... Proficiency in model quantization techniques (FP8, AWQ, GPTQ) and performance tuning Experience ...
... boosting, deep learning, graph-based models). Technology Familiarity * Supervised learning ... Model serving and inference optimization (vLLM-class serving, quantization). Nice-to-Have:
New
Quick apply
... boosting, deep learning, graph-based models). Technology Familiarity * Supervised learning ... Model serving and inference optimization (vLLM-class serving, quantization). Nice-to-Have:
New
AI Engineer
Charlotte, NC · On-site
$50K - $112K/yr
... networks and deep learning methods for advanced AI applications - Managing data quality and ... using quantization, inference acceleration, and model-routing techniques - Designing agent ...
AI Engineer
Charlotte, NC · On-site
$50K - $112K/yr
... networks and deep learning methods for advanced AI applications - Managing data quality and ... using quantization, inference acceleration, and model-routing techniques - Designing agent ...
Deep Learning Quantization information
See Charlotte, NC salary details
$21.3K is the 25th percentile. Wages below this are outliers.
$10.7K - $22.2K
27% of jobs
$22.2K - $33.7K
0% of jobs
$33.7K - $45.1K
0% of jobs
$45.1K - $56.6K
0% of jobs
$56.6K - $68K
0% of jobs
The median wage is $78.5K / yr.
$68K - $79.5K
25% of jobs
$79.5K - $90.9K
18% of jobs
$99.1K is the 75th percentile. Wages above this are outliers.
$90.9K - $102.4K
7% of jobs
$102.4K - $113.8K
2% of jobs
$113.8K - $125.3K
0% of jobs
$125.3K - $136.7K
21% of jobs
$10.7K
$81.9K
$136.7K
How much do deep learning quantization jobs pay per year?
What are the key skills and qualifications needed to thrive as a deep learning quantization engineer, and why are they important?
What is the difference between Deep Learning Quantization vs Machine Learning Engineer?
| Aspect | Deep Learning Quantization | Machine Learning Engineer |
|---|---|---|
| Required Credentials | Advanced degrees in AI, Computer Science, or related fields; knowledge of neural networks | Bachelor's or Master's in CS, Data Science, or related fields; programming skills |
| Work Environment | Research labs, AI development teams, hardware optimization settings | Software development teams, data-driven projects, product-focused environments |
| Industry Usage | AI hardware optimization, model deployment, edge computing | Model development, data analysis, software solutions across industries |
Deep Learning Quantization focuses on reducing model size and improving inference speed through techniques like weight and activation quantization, often in hardware or embedded systems. Machine Learning Engineers develop, implement, and optimize machine learning models for various applications. While both roles require knowledge of AI and programming, Deep Learning Quantization is more specialized in model optimization techniques, whereas Machine Learning Engineers work broadly on model development and deployment.
What is deep learning quantization?
What are some common challenges faced when implementing deep learning quantization in production environments?
- Junior Machine Learning Engineer
- Urgently Hiring Machine Learning Engineer New Grad
- Hourly Remote Machine Learning Engineer
- Remote Senior Machine Learning Engineer
- Senior Machine Learning Researcher
- Senior Machine Learning Engineer
- Machine Learning Engineer Apprenticeship
- Google Cloud Machine Learning Engineer
- Machine Learning Engineer Intern
- Remote Machine Learning Engineer
Full-time
Re-posted 22 days ago
Infosys rating
7.0
Based on 62 frontline employees who took The Breakroom Quiz
149th of 221 rated it services
Job description
Infosys is a global leader in next-generation digital services and consulting, specializing in transforming data into actionable insights. As a Data Science Consultant 2, you will develop and optimize predictive models and algorithms, ensuring data readiness for advanced modeling and contributing to innovative analytics solutions.
Responsibilities:
• Develop data preparation tasks, while identifying patterns or anomalies.
• Ensure data readiness for advanced modeling.
• Develop models for complex use cases (e.g., forecasting models, LLM-based solutions), while refining algorithms to meet business needs, and ensure smooth deployment into scalable, production-ready solutions.
• Conduct testing and optimize algorithms for performance, reliability, and scalability, while providing guidance to team members in best practices.
• Design and develop predictive models and data-driven analyses to address business challenges.
• Build, evaluate, and deploy models, standardize code, and contribute to knowledge management.
• Leverage tools like SAS and R/Python to create reusable customizations for non-ML, ML, and deep learning algorithms, while enhancing analytics including LLMs, and create innovative, cost-effective solutions.
• Define analytics problems for projects; execute visualization, analysis, and predictive modeling under guidance.
• Proactively maintain models and implement improvements for accuracy and reliability.
• Apply governance controls to mitigate risks and ensure compliance.
• Analyze performance trends, recommend improvements, and document discrepancies for escalation.
• Maintain comprehensive documentation standards, while participating in knowledge transfer sessions.
• Participate in discussions with stakeholders to refine requirements, provide insights, and guide implementation of models.
• Apply the predefined quality measurement framework at an individual task level in the project.
• Deploy complex analytics tools or multi-system integration, while validating deployment success.
• Participate in developing scripts or templates for repeated deployments tasks.
• Contribute to analytic solutions, IP asset creation, and training initiatives.
• Contribute to thought leadership such as papers, innovative non-ML, ML, deep learning or LLM models, and proofs of concepts.
• Participate in and deliver analytics training, while contributing to content creation.
• Provide input for segment and unit-level business plans.
Qualifications:
Required:
• vLLM, TensorRT-LLM, Triton, SGLang
• Quantization (FP8/AWQ/GPTQ), tensor parallelism
• Performance benchmarking & tuning
• Kubernetes, GKE, KServe / ML serving patterns
• Helm, Operators
• GPU orchestration concepts and scheduling patterns
• GCP and/or Azure (strong hands-on)
• Terraform
• Cloud networking, landing zones, governance/org policies
• HashiCorp Vault (secrets management)
• Observability & SRE
• Prometheus/Grafana, logging, tracing
• SRE/SLO mindset, reliability engineering
• Bachelor’s degree or foreign equivalent required from an accredited institution. Will also consider three years of progressive experience in the specialty in lieu of every year of education.
• This position may require relocation and/or travel to work/project location.
• Candidates authorized to work for any employer in the United States without employer-based visa sponsorship are welcome to apply. Infosys is unable to provide immigration sponsorship for this role now or in the future.
Preferred:
• Experience in Big Data technologies (e.g., BigQuery, Hadoop).
• Expertise in ML model development, data engineering, and software engineering principles.
• Knowledge of MLOps and AI/ML deployment (e.g., SageMaker, Snowflake). Familiarity with CI/CD, DevOps, and automation tools in AI/ML contexts.
• Design and implement LLM inference serving stacks using: vLLM, TensorRT-LLM, Triton Inference Server, SGLang
• Inference optimization techniques: continuous batching, speculative decoding, KV/prefix caching
• Quantization: FP8 / AWQ / GPTQ and tuning for GPU utilization
• Build Kubernetes-based serving platforms: KServe, Kubernetes ML Serving, GKE, OpenShift (OCP) (where applicable)
• Enable GenAI platforms and RAG use cases: Integrate LLM services with RAG pipelines, Provide reusable internal libraries, templates, and developer enablement assets
• Collaborate with cross-functional teams and client stakeholders to productionize LLM workloads at scale
Company:
Infosys is a technology company that offers consulting, outsourcing, cloud infrastructure, program management, and software services. Founded in 1981, the company is headquartered in Bangalore, IND, with a team of 10001+ employees. The company is currently Late Stage.
About Infosys
Sourced by ZipRecruiter
Industry
Business management consulting
Company size
10,000+ Employees
Headquarters location
Bengaluru, KA, IN