1

Ml Inference Jobs in Newark, NJ (NOW HIRING)

Software Engineer - Infrastructure

New York, NY ยท On-site +1

$165K - $330K/yr

Develop infrastructure components for our ML inference platform using Python and Go * Implement and maintain Kubernetes deployments for model serving * Contribute to our inference orchestration layer ...

Senior ML Infrastructure Engineer

New York, NY ยท On-site

$118K - $161K/yr

Hands-on experience with managed ML inference and serving platforms such as AWS SageMaker and GCP Vertex AI. * A proven track record operating inference at large scale across a range of model types ...

Senior Site Reliability Engineer

New York, NY ยท On-site

$62.25 - $82.75/hr

Build and operate ML inference infrastructure - model serving, GPU workloads, language model gateway and routing - contributing to a broader systems portfolio * Help set infrastructure standards and ...

Software Engineer, AI/ML

New York, NY ยท On-site

$165K - $225K/yr

Can be anywhere in that lifecycle, from training machine learning systems, to creating the interfaces users use to navigate inference results. * An interest and passion for AI/ML systems, if you ...

Software Engineer, AI/ML

New York, NY ยท On-site

$165K - $225K/yr

Can be anywhere in that lifecycle, from training machine learning systems, to creating the interfaces users use to navigate inference results. * An interest and passion for AI/ML systems, if you ...

Implement, refine, and productionize cutting-edge techniques (quantization, speculative decoding, KV cache reuse, chunked prefill and LoRA) for ML model inference and infrastructure. * Deep dive into ...

Technical Program Manager, Inference

New York, NY ยท On-site

$141K - $182K/yr

The AI/ML TPM team owns delivery and execution across CoreWeave's AI/ML Platform Services ... The Inference team is responsible for building and operating highly scalable, reliable production ...

Technical Program Manager, Inference

Livingston, NJ ยท On-site

$140K - $182K/yr

The AI/ML TPM team owns delivery and execution across CoreWeave's AI/ML Platform Services ... The Inference team is responsible for building and operating highly scalable, reliable production ...

next page

Showing results 1-20

Ml Inference information

See Newark, NJ salary details

$39.2K

$128.4K

$205.5K

How much do ml inference jobs pay per year?

As of Sep 8, 2026, the average yearly pay for ml inference in Newark, NJ is $128,351.00, according to ZipRecruiter salary data. Most workers in this role earn between $103,000.00 and $142,200.00 per year, depending on experience, location, and employer.

What is ML inference?

ML inference refers to the process of using a trained machine learning model to make predictions or decisions based on new data. After a model has been trained on historical data, inference is the phase where that model is deployed and used in real-world applications, such as recognizing speech, detecting objects in images, or recommending products. The focus in ML inference is on speed, efficiency, and scalability to ensure quick predictions, often in real time. This process is critical for practical applications like mobile apps, web services, and embedded systems. Optimizing inference involves reducing latency, memory usage, and computational requirements.

What are the key skills and qualifications needed to thrive in ML inference?

To thrive in ML Inference, you need a solid background in machine learning principles, programming (Python or C++), and experience with deploying models at scale, often supported by a degree in computer science or a related field. Familiarity with frameworks and tools such as TensorFlow, PyTorch, ONNX, and cloud platforms like AWS SageMaker or Google AI Platform is typically required. Strong problem-solving skills, attention to detail, and effective communication are crucial soft skills for collaborating with multidisciplinary teams and optimizing model performance. These skills ensure efficient, scalable, and reliable deployment of machine learning solutions in real-world applications.

What are some common challenges faced by ML inference engineers when deploying models to production?

ML Inference Engineers often encounter challenges such as optimizing model latency and throughput to meet production requirements, ensuring compatibility with diverse hardware environments, and managing model versioning and updates without disrupting service. Additionally, balancing resource utilization and inference accuracy while monitoring real-time performance metrics is crucial. Collaboration with data scientists, DevOps, and software engineers is typically essential to streamline deployment and maintain robust, scalable inference pipelines.

What is the difference between Ml Inference vs Data Scientist?

AspectML InferenceData Scientist
Required CredentialsKnowledge of machine learning models, programming skillsDegree in data science, statistics, or related fields
Work EnvironmentDeploying models in production, real-time data processingData analysis, model development, research
Industry UsageAI product deployment, software companiesResearch institutions, tech firms, consulting

ML Inference focuses on deploying trained models to make predictions on new data, often in real-time. Data Scientists develop and analyze models, working primarily in research and development. While both roles require understanding of machine learning, ML Inference emphasizes deployment and operationalization, whereas Data Scientists focus on model creation and analysis.

What job categories do people searching Ml Inference jobs in Newark, NJ look for?

The top searched job categories for Ml Inference jobs in Newark, NJ are:

What cities near Newark, NJ are hiring for Ml Inference jobs?

Cities near Newark, NJ with the most Ml Inference job openings:

Infographic showing various Ml Inference job openings in Newark, NJ as of July 2026, with employment types broken down into 96% Full Time, 2% Part Time, and 2% Contract. Highlights an 81% Physical, 5% Hybrid, and 14% Remote job distribution, with an average salary of $128,351 per year, or $61.7 per hour.

Machine Learning Performance Engineer (Inference)

Tower Research Capital

New York, NY โ€ข On-site

$200K - $300K/yr

Full-time

PTO

Posted 27 days ago


Job description

Tower Research Capital is a leading quantitative trading firm founded in 1998. Tower has built its business on a high-performance platform and independent trading teams. We have a 25+ year track record of innovation and a reputation for discovering unique market opportunities.
Tower is home to some of the world's best systematic trading and engineering talent. We empower portfolio managers to build their teams and strategies independently while providing the economies of scale that come from a large, global organization.ย 

Engineers thrive at Tower while developing electronic trading infrastructure at a world class level. Our engineers solve challenging problems in the realms of low-latency programming, FPGA technology, hardware acceleration and machine learning. Our ongoing investment in top engineering talent and technology ensures our platform remains unmatched in terms of functionality, scalability and performance.

At Tower, every employee plays a role in our success. Our Business Support teams are essential to building and maintaining the platform that powers everything we do - combining market access, data, compute, and research infrastructure with risk management, compliance, and a full suite of business services. Our Business Support teams enable our trading and engineering teams to perform at their best.

At Tower, employees will find a stimulating, results-oriented environment where highly intelligent and motivated colleagues inspire each other to reach their greatest potential.

Summary:
As part of Tower Research's Core Engineering team, you will bridge the gap between quantitative research and high-performance production systems, architecting inference pipelines that operate at the physical limits of hardware. Your objective will be to drive the speed, efficiency, and reliability of our ML inference pipelines to their absolute limits, ensuring our predictive models consistently achieve microsecond-level latency.
Responsibilities:

  • Benchmarking & Strategy:ย 
    • Lead the technical evaluation of diverse inference platforms - ranging across CPUs, GPUs, and FPGAs - to guide Tower's infrastructure deployment decisions.
  • System Architecture Optimization:ย 
    • Analyze and enhance execution across deep memory hierarchies to maximize resource utilization and parallel processing. You will assess and resolve memory subsystem and interconnect bottlenecks across the end-to-end inference lifecycle.
  • Infrastructure & Deployment Feasibility:ย 
    • Collaborate with Infrastructure teams to understand thermal, power, and operational constraints of hardware platforms to design inference strategies for our latency-critical trading strategies that fit within those envelopes.
  • GPU Kernel Development:ย 
    • Develop highly optimized kernels and integrate specialized performance libraries to extract maximum computational throughput from the underlying silicon.
  • Model Optimization & Deployment:ย 
    • Implement advanced model reduction techniques (quantization, pruning, distillation) to ensure compact memory footprints and numerical stability. Prioritize optimization for low-latency, event-level inference workloads to meet real-time trading requirements.
  • Cross-Functional Collaboration:ย 
    • Collaborate closely with ML Researchers, HPC Engineers, FPGA Engineers, and Datacenter Engineers to bring to fruition target deployments.

Qualifications:

  • 2+ years of experience optimizing deep learning inference in latency-sensitive or high-throughput production environments, in any domain.
  • ML Frameworks: Deep expertise in lower-level ML framework development (PyTorch/JAX), paired with strong Python/C++ skills and a thorough understanding of mixed-precision computation.
  • Kernel Development & Optimization Tooling: Proven experience in custom GPU kernel development. Deep familiarity with advanced optimization libraries and compilers (e.g., Triton, TensorRT, ONNX, IREE, HLS4ML, cuBLAS, CUTLASS) as well as profiling tools (e.g., Nsight Systems, Nsight Compute).
  • GPU Architecture Mastery: Deep expertise in GPU microarchitecture, encompassing SM execution, warp scheduling, and full memory hierarchy optimization (registers to HBM).
  • Cross-Architecture Benchmarking: Proven record of rigorous, data-driven approach to evaluating inference performance across heterogeneous compute architectures.
  • Bonus: Practical experience targeting and optimizing inference workloads on specialized hardware ecosystems, including FPGAs and ASICs.
  • Prior experience in financial trading is not required.


Anticipated annual base salary range $200,000-$300,000, plus eligible for discretionary bonus.

Tower's headquarters are in the historic Equitable Building, right in the heart of NYC's Financial District and our impact is global, with over a dozen offices around the world.

ย At Tower, we believe work should be both challenging and enjoyable. That is why we foster a culture where smart, driven people thrive - without the egos. Our open concept workplace, casual dress code, and well-stocked kitchens reflect the value we place on a friendly, collaborative environment where everyone is respected, and great ideas win.

Our benefits include:

  • Generous paid time off policies
  • Savings plans and other financial wellness tools available in each region
  • Hybrid working opportunities
  • Free breakfast, lunch, and snacks daily
  • In-office wellness experiences and reimbursement for select wellness expenses (e.g., gym, personal training and more)
  • Company-sponsored sports teams and fitness events (JPM Corporate Challenge, Cycle for Survival, Wall Street Rides FAR and more)
  • Volunteer opportunities and charitable giving
  • Social events, happy hours, treats, and celebrations throughout the year
  • Workshops and continuous learning opportunities

At Tower, you'll find a collaborative and welcoming culture, a diverse team and a workplace that values both performance and enjoyment. No unnecessary hierarchy. No ego. Just great people doing great work - together.

Tower Research Capital is an equal opportunity employer.