1

Ml Inference Jobs in New York (NOW HIRING)

ML Ops Engineer Location: Iselin NJ Experience Required Experience building production Al/ML systems at scale Deploying real-time ML inference pipelines processing millions of records at high ...

Deploying real-time ML inference pipelines processing millions of records at high throughput * Experience building end to end automated MLOps capabilities along with model and feature drift ...

Solution Architect (AI/LLM Inference)

New York, NY ยท On-site +1

$165K - $330K/yr

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies ... AI/ML background and the ability to credibly discuss AI/ML topics with technical stakeholders.

Software Engineer - Infrastructure

New York, NY ยท On-site +1

$165K - $330K/yr

Develop infrastructure components for our ML inference platform using Python and Go * Implement and maintain Kubernetes deployments for model serving * Contribute to our inference orchestration layer ...

Machine Learning Engineer

New York, NY ยท Hybrid

$145K - $180K/yr

The ML Engineer has hands-on experience building and optimizing ML inference systems that run in production environments. This role will develop and tune pipelines that transform millions of photos ...

Machine Learning Engineer

Manhattan, NY ยท Hybrid

$145K - $180K/yr

The ML Engineer has hands-on experience building and optimizing ML inference systems that run in production environments. This role will develop and tune pipelines that transform millions of photos ...

Machine Learning Engineer

Manhattan, NY ยท On-site

$145K - $180K/yr

The ML Engineer has hands-on experience building and optimizing ML inference systems that run in production environments. This role will develop and tune pipelines that transform millions of photos ...

Senior ML Infrastructure Engineer

New York, NY ยท On-site

$118K - $161K/yr

Hands-on experience with managed ML inference and serving platforms such as AWS SageMaker and GCP Vertex AI. * A proven track record operating inference at large scale across a range of model types ...

Work with model developers to tune their neural networks for better inference efficiency and ... working with ML inference or linear algebra computation * C++ programming skills, including ...

Software Engineer, AI/ML

New York, NY ยท On-site

$165K - $225K/yr

Can be anywhere in that lifecycle, from training machine learning systems, to creating the interfaces users use to navigate inference results. * An interest and passion for AI/ML systems, if you ...

Software Engineer, AI/ML

New York, NY ยท On-site

$165K - $225K/yr

Can be anywhere in that lifecycle, from training machine learning systems, to creating the interfaces users use to navigate inference results. * An interest and passion for AI/ML systems, if you ...

Implement, refine, and productionize cutting-edge techniques (quantization, speculative decoding, kv cache reuse, chunked prefill and LoRA) for ML model inference and infrastructure. * Deep dive into ...

Technical Program Manager, Inference

Livingston, NJ ยท On-site

$140K - $182K/yr

The AI/ML TPM team owns delivery and execution across CoreWeave's AI/ML Platform Services ... The Inference team is responsible for building and operating highly scalable, reliable production ...

next page

Showing results 1-20

Ml Inference information

What is a $900000 AI job?

A $900,000 AI job typically refers to high-level roles in artificial intelligence, such as senior machine learning engineers or AI research directors, often involving advanced skills in deep learning, data modeling, and programming with tools like Python and TensorFlow. These positions usually require extensive experience, specialized knowledge, and may include leadership responsibilities or strategic decision-making.

What is ML inference?

ML inference refers to the process of using a trained machine learning model to make predictions or decisions based on new data. After a model has been trained on historical data, inference is the phase where that model is deployed and used in real-world applications, such as recognizing speech, detecting objects in images, or recommending products. The focus in ML inference is on speed, efficiency, and scalability to ensure quick predictions, often in real time. This process is critical for practical applications like mobile apps, web services, and embedded systems. Optimizing inference involves reducing latency, memory usage, and computational requirements.

What is the difference between Ml Inference vs Data Scientist?

AspectML InferenceData Scientist
Required CredentialsKnowledge of machine learning models, programming skillsDegree in data science, statistics, or related fields
Work EnvironmentDeploying models in production, real-time data processingData analysis, model development, research
Industry UsageAI product deployment, software companiesResearch institutions, tech firms, consulting

ML Inference focuses on deploying trained models to make predictions on new data, often in real-time. Data Scientists develop and analyze models, working primarily in research and development. While both roles require understanding of machine learning, ML Inference emphasizes deployment and operationalization, whereas Data Scientists focus on model creation and analysis.

What engineer makes $500,000 a year?

Senior machine learning engineers with extensive experience, advanced skills in deep learning, and expertise in deploying large-scale models can earn salaries approaching or exceeding $500,000 annually, especially in high-cost-of-living areas or top tech companies. Compensation often includes base salary, bonuses, and stock options, reflecting their specialized knowledge and impact on product development.

Which 3 jobs will survive AI?

Jobs involving Ml Inference, such as data scientists, machine learning engineers, and AI system architects, are likely to persist as they require specialized expertise in developing, deploying, and maintaining AI models. These roles demand critical thinking, domain knowledge, and skills in programming and data analysis that are less easily automated. Continuous learning and staying updated with AI tools and frameworks are essential for these professions to remain relevant.

What are some common challenges faced by ML Inference Engineers when deploying models to production?

ML Inference Engineers often encounter challenges such as optimizing model latency and throughput to meet production requirements, ensuring compatibility with diverse hardware environments, and managing model versioning and updates without disrupting service. Additionally, balancing resource utilization and inference accuracy while monitoring real-time performance metrics is crucial. Collaboration with data scientists, DevOps, and software engineers is typically essential to streamline deployment and maintain robust, scalable inference pipelines.

Will MLE be replaced by AI?

Machine Learning Engineers (MLEs) design, develop, and optimize AI models and systems. While AI automation tools can assist with certain tasks, MLEs are essential for building, tuning, and maintaining complex models, making complete replacement unlikely in the near term. Their expertise in data handling, model deployment, and system integration remains critical in AI development environments.

What are the key skills and qualifications needed to thrive in ML Inference, and why are they important?

To thrive in ML Inference, you need a solid background in machine learning principles, programming (Python or C++), and experience with deploying models at scale, often supported by a degree in computer science or a related field. Familiarity with frameworks and tools such as TensorFlow, PyTorch, ONNX, and cloud platforms like AWS SageMaker or Google AI Platform is typically required. Strong problem-solving skills, attention to detail, and effective communication are crucial soft skills for collaborating with multidisciplinary teams and optimizing model performance. These skills ensure efficient, scalable, and reliable deployment of machine learning solutions in real-world applications.
What cities in New York are hiring for Ml Inference jobs? Cities in New York with the most Ml Inference job openings:
Senior ML Engineer (GCP) - Remote / Telecommute

Senior ML Engineer (GCP) - Remote / Telecommute

CYNET SYSTEMS

New York, NY โ€ข Remote

$65 - $70/hr

Contractor

Posted 15 days ago


Job description

Job Overview:

Pay Range:ย $65hr - $70hr

Requirement/Must Have:

  • 10+ years of experience.
  • GCP cloud experience is mandatory.
  • Strong foundation in ML inference, deployment, and quality testing.
  • Demonstrated ability to ramp up quickly on new and unfamiliar tech stacks.
  • End-to-end problem-solving mindset.
  • Core ML knowledge sufficient to benchmark models and collaborate with researchers.
  • Experience deploying models in cloud environments, ideally GCP.

Responsibilities:

  • Evaluate and benchmark new ML inference frameworks to guide production decisions.
  • Deploy models to GCP and integrate them into production applications and Java-based streaming pipelines.
  • Own deployment automation end-to-end โ€” from model handoff through live serving.
  • Monitor how models behave in production for real end-users.
  • Design and execute benchmarking, performance testing, and quality testing on ML models.
  • Perform model sampling to support quality evaluation and researcher feedback loops.
  • Debug issues across the full stack โ€” from inference layer down to streaming pipelines.
  • Partner with ML researchers to provide benchmarking feedback and guide inference decisions.
  • Adapt rapidly to non-standard and evolving tech stacks across hybrid (on-prem + GCP) infrastructure.

Nice to Have:

  • Exposure to Java or JVM-based systems.
  • Familiarity with streaming data architectures.
  • Experience in hybrid cloud/on-prem environments.

Skills:

  • Machine Learning frameworks: TensorFlow, PyTorch, JAX or similar.
  • Primary platform: Google Cloud Platform (inference, deployment automation, experimentation, sampling).
  • Production integration: Java-based streaming pipelines.
  • Infrastructure: Hybrid โ€” on-premise streaming + GCP serving stacks.
  • Distributed systems: Working knowledge required for debugging and end-to-end testing.

Qualification And Education:

  • Bachelor's or Masterโ€™s degree in Computer Science, Computer or Electrical Engineering, Mathematics, or a related field.

Founded in 2010 and headquartered in the Washington, DC metro area, Cynet Systems Inc. is a leading staffing and recruiting powerhouse. Proudly recognized as a nationally and locally certified diversity firm, Cynet delivers agile, scalable talent solutions across industries. With an active footprint in all 50 U.S. states and Canada, we support thousands of consultants through our expansive, high-performing recruitment engine operating across North America and Asiaโ€”ensuring speed, quality, and consistency in every hire.

Cynet Systems logo

About Cynet Systems

Sourced by ZipRecruiter

Cynet Systems Inc is a staffing and recruiting corporation nestled in Ashburn, VA, USA. Established in 2010, the company operates within the Information Technology and Services sector, specializing in providing effective workforce solutions to different business needs, including IT consulting, direct hire, and contract staffing services. Through the years, Cynet Systems has built an impressive portfolio, going beyond borders and expanding its operations internationally in Canada and India. Rooted in its core values of teamwork, leadership, and commitment, Cynet Systems helps businesses unlock their full potential by providing versatile and competent professionals that perfectly align with their needs. Fueled by their unwavering mission to deliver top-tier talent to businesses worldwide, Cynet Systems garnered various recognitions including SIA's fastest-growing staffing firms and Best Place to Work in Virginia for 2019.

Industry

It services

Company size

501 - 1,000 Employees

Headquarters location

Sterling, VA, US

Year founded

2010

Social media