1

Deep Learning Quantization Jobs in Bethpage, NY (NOW HIRING)

Implement techniques such as distillation, quantization, and pruning to aggressively accelerate ... Strong experience in deep learning systems and infrastructure * Expertise in PyTorch, CUDA, Triton ...

... quantization, compression, and resource-efficient AI, to drive performance improvements and ... Research experience in machine learning, deep learning, natural language processing, and/or ...

AI Engineer

Melville, NY ยท On-site

$50K - $112K/yr

... networks and deep learning methods for advanced AI applications - Managing data quality and ... using quantization, inference acceleration, and model-routing techniques - Designing agent ...

AI Engineer

Stamford, CT ยท On-site

$50K - $112K/yr

... networks and deep learning methods for advanced AI applications - Managing data quality and ... using quantization, inference acceleration, and model-routing techniques - Designing agent ...

AI Engineer

New York, NY ยท On-site

$50K - $112K/yr

... networks and deep learning methods for advanced AI applications - Managing data quality and ... using quantization, inference acceleration, and model-routing techniques - Designing agent ...

Deep dive into underlying codebases of TensorRT, PyTorch, TensorRT-LLM, vllm, sglang, CUDA, and ... Familiarity with LLM optimization techniques (e.g., quantization, speculative decoding, continuous ...

Sr. AI Engineer

Manhattan, NY ยท Remote

$114K - $157K/yr

Optimize inference performance and cost efficiency through techniques such as model quantization ... learning, and deep learning 5. Experience with AI platforms like PyTorch or TensorFlow 6. ...

Sr. AI Engineer

New York, NY ยท On-site

$114K - $157K/yr

Optimize inference performance and cost efficiency through techniques such as model quantization ... learning, and deep learning 5. Experience with AI platforms like PyTorch or TensorFlow 6. ...

AI Researcher

New York, NY ยท On-site

$175K - $250K/yr

... data analysis, vector quantization, decision tree methods, EM methods, Bayesian methods ... Demonstration of deep knowledge of large language models and deep neural networks for practical ...

AI Researcher

New York, NY ยท On-site

$175K - $250K/yr

... data analysis, vector quantization, decision tree methods, EM methods, Bayesian methods ... Demonstration of deep knowledge of large language models and deep neural networks for practical ...

AI Researcher - Vatic Labs

Manhattan, NY ยท On-site

$175K - $250K/yr

... data analysis, vector quantization, decision tree methods, EM methods, Bayesian methods ... Demonstration of deep knowledge of large language models and deep neural networks for practical ...

Showing results 21-40

Deep Learning Quantization information

See Bethpage, NY salary details

$11.3K

$85.9K

$143.3K

How much do deep learning quantization jobs pay per year?

As of Aug 7, 2026, the average yearly pay for deep learning quantization in Bethpage, NY is $85,879.00, according to ZipRecruiter salary data. Most workers in this role earn between $73,700.00 and $142,300.00 per year, depending on experience, location, and employer.

What are the key skills and qualifications needed to thrive as a deep learning quantization engineer, and why are they important?

To excel as a Deep Learning Quantization Engineer, you need a strong background in machine learning, applied mathematics, and computer science, usually supported by an advanced degree in a related field. Familiarity with deep learning frameworks (such as TensorFlow or PyTorch), quantization toolkits, and hardware acceleration platforms is crucial. Analytical thinking, problem-solving, and clear technical communication are standout soft skills in this role. These abilities are essential for efficiently optimizing models for deployment on resource-constrained hardware while maintaining accuracy and performance.

What is the difference between Deep Learning Quantization vs Machine Learning Engineer?

AspectDeep Learning QuantizationMachine Learning Engineer
Required CredentialsAdvanced degrees in AI, Computer Science, or related fields; knowledge of neural networksBachelor's or Master's in CS, Data Science, or related fields; programming skills
Work EnvironmentResearch labs, AI development teams, hardware optimization settingsSoftware development teams, data-driven projects, product-focused environments
Industry UsageAI hardware optimization, model deployment, edge computingModel development, data analysis, software solutions across industries

Deep Learning Quantization focuses on reducing model size and improving inference speed through techniques like weight and activation quantization, often in hardware or embedded systems. Machine Learning Engineers develop, implement, and optimize machine learning models for various applications. While both roles require knowledge of AI and programming, Deep Learning Quantization is more specialized in model optimization techniques, whereas Machine Learning Engineers work broadly on model development and deployment.

What is deep learning quantization?

Deep learning quantization is the process of reducing the precision of the numbers used to represent a neural network's parameters, activations, or both. By converting the typically used 32-bit floating-point values to lower bit-width formats such as 16-bit or 8-bit integers, quantization significantly reduces the memory footprint and computational requirements of deep learning models. This technique helps deploy models efficiently on edge devices and mobile hardware while maintaining acceptable accuracy levels. Quantization is widely used in model optimization for faster inference and lower power consumption.

What are some common challenges faced when implementing deep learning quantization in production environments?

One of the main challenges in implementing deep learning quantization is balancing model accuracy with computational efficiency, as quantization can sometimes lead to a drop in model performance. Additionally, ensuring hardware compatibility and optimizing for different devices (such as CPUs, GPUs, or edge devices) can require extensive testing and tuning. Collaboration with data scientists, software engineers, and hardware specialists is often essential to successfully deploy quantized models at scale. Staying updated with the latest quantization techniques and frameworks is also important for overcoming these challenges.
What cities near Bethpage, NY are hiring for Deep Learning Quantization jobs? Cities near Bethpage, NY with the most Deep Learning Quantization job openings:
Infographic showing various Deep Learning Quantization job openings in Bethpage, NY as of June 2026, with employment types broken down into 2% As Needed, 25% Full Time, 68% Part Time, and 5% Contract. Highlights an 87% Physical, 2% Hybrid, and 11% Remote job distribution, with an average salary of $85,879 per year, or $41.3 per hour.

Research Engineer, Knowledge Graph Intelligence

Point72

New York, NY โ€ข On-site

$175K - $250K/yr

Full-time

Retirement

Re-posted 12 days ago


Job description

A Career with point72's Surveillance team
Point72's Surveillance team sets the industry standard for intelligence-driven surveillance by proactively identifying, monitoring, and assessing various sources of compliance risk using proprietary tools and specialized tradecraft. We support senior management by providing strategic assessments, actionable recommendations, and real-time escalations. At Point72, members of the Surveillance team conduct integrated trade and communication surveillance and collaborate to turn information into intelligence for our internal customers. The team also monitors employee activity for evidence of violations of applicable federal securities laws, internal compliance policies and procedures, and relevant rules and regulations enforced by the SEC, FINRA, and other organizations.
What you'll do
As a Machine Learning Engineer - Applied Scientist you will play a critical role in developing algorithmic solutions and models for production-ready applications that support our front office investment professionals. You will specialize in natural language processing (NLP) solutions that extract insights from unstructured text data, with additional capabilities in predictive modeling, clustering, and time series analysis. You will manage all aspects of the research process including methodology selection, data collection and analysis, implementation and testing, prototyping, and performance evaluation. You will apply, adapt, and extend existing results in the broad field of NLP, while also conducting novel research as required. Specifically, you will:
  • Contribute to projects across various machine learning (ML) disciplines, including NLP, unstructured data analysis, predictive modeling, and classic machine learning.
  • Implement GenAI solutions, utilize ML infrastructure, and contribute to modeling, data preparation, optimization, and performance enhancements.
  • Work with sparse data and apply techniques to improve model accuracy and generalization.
  • Conduct data evaluation, including data preprocessing, feature engineering, and model performance assessment.
  • Collaborate cross-functionally with data engineers, software developers, and product teams to integrate models into production systems.
  • Stay up to date with the latest advancements in natural language processing and machine learning, applying new techniques as needed.

What's REQUIRED
  • PhD, master's degree, or 4+ years of CS, CE, ML or related field experience.
  • 6+ years of experience building ML models and developing algorithms.
  • Strong proficiency in Python, and hands-on experience with NumPy, Hugging Face, PyTorch, and spaCy for NLP applications.
  • Prior experience in the domains of LLMs, foundation models, or large-scale deep learning systems, with a complete understanding of modern training, fine-tuning, quantization, and model evaluation.
  • Expertise in working with sparse data and applying techniques such as data augmentation, weak supervision, and semi-supervised learning.
  • Solid grasp of NLP concepts, including tokenization, embeddings, attention mechanisms, and transformer-based architectures.
  • Experience with data evaluation techniques, model explainability, and error analysis.
  • Experience working in a Linux environment.
  • Commitment to the highest ethical standards.

We take care of our people
We invest in our people, their careers, their health, and their well-being. When you work here, we provide:
  • Fully-paid health care benefits
  • Generous parental and family leave policies
  • Volunteer opportunities
  • Support for employee-led affinity groups representing women, people of color, and the LGBT+ community
  • Mental and physical wellness programs
  • Tuition assistance
  • A 401(k) savings program with an employer match and more

About Point72
Point72 is a leading global alternative investment firm led by Steven A. Cohen. Building on more than 30 years of investing experience, Point72 seeks to deliver superior returns for its investors through fundamental and systematic investing strategies across asset classes and geographies. We aim to attract and retain the industry's brightest talent by cultivating an investor-led culture and committing to our people's long-term growth. For more information, visit https://point72.com/.
The annual base salary range for this role is $175,000-$250,000 (USD) , which does not include discretionary bonus compensation or our comprehensive benefits package. Actual compensation offered to the successful candidate may vary from posted hiring range based upon geographic location, work experience, education, and/or skill level, among other things.