1

Ml Inference Jobs in Newark, NJ (NOW HIRING)

Solution Architect

Manhattan, NY · On-site

$120 - $150/hr

Experience running or supporting benchmarks for ML inference deployments. * Familiarity with infrastructure tradeoffs relevant to inference performance and cost (for example GPU selection and latency ...

Senior ML Ops Engineer

Manhattan, NY · On-site

$115K - $158K/yr

Responsibilities : • Own ML pipelines end to end -- experimentation to production -- and the infrastructure behind training, inference, and agentic workloads • Give the AI/ML team a paved road ...

Solutions Architect

New York, NY · On-site +1

$165K - $330K/yr

Experience running or supporting benchmarks for ML inference deployments. * Familiarity with infrastructure tradeoffs relevant to inference performance and cost (for example GPU selection and latency ...

Significant experience building production ML systems, inference runtimes, or performance-sensitive infrastructure * Strong programming ability in Python and C++ * Deep understanding of transformer ...

Manager, Solutions Architect

New York, NY · On-site

$165K - $330K/yr

Experience running or supporting benchmarks for ML inference deployments. * Familiarity with infrastructure tradeoffs relevant to inference performance and cost (for example GPU selection and latency ...

Senior Backend Engineer (ML)

New York, NY · On-site

$200K - $230K/yr

Architect, build, and scale the cloud infrastructure behind our ML training, inference, and data ... pipelines. * Design resilient systems for model deployment, evaluation, and monitoring that stay ...

Technical Program Manager, Inference

Manhattan, NY · On-site

$142K - $183K/yr

... ML platform engineering • Proven experience driving large-scale infrastructure or platform ... distributed inference systems, GPU compute, cloud-native architectures, and performance ...

Lead DevOps Engineer

Manhattan, NY · On-site

$58 - $79.50/hr

Key Projects - Build and optimize real-time serving infrastructure for personalization and engagement (including ML-inference workloads). - Develop scalable, secure CI/CD pipelines for deploying ...

Silicon Software Lead

New York, NY · On-site

$240K - $310K/yr

Deep understanding of ML inference workloads and the constraints that shape their execution on hardware * Experience with compiler frameworks such as MLIR or LLVM, or with inference runtimes and ...

Develop causal inference methodologies to understand true incrementality of product changes ... Proven track record building and deploying ML models in production , particularly in ...

Lead DevOps Engineer

Manhattan, NY · On-site

$58 - $79.50/hr

Key Projects - Build and optimize real-time serving infrastructure for personalization and engagement (including ML-inference workloads). - Develop scalable, secure CI/CD pipelines for deploying ...

Showing results 41-60

Ml Inference information

See Newark, NJ salary details

$39.2K

$128.4K

$205.5K

How much do ml inference jobs pay per year?

As of Aug 11, 2026, the average yearly pay for ml inference in Newark, NJ is $128,351.00, according to ZipRecruiter salary data. Most workers in this role earn between $103,000.00 and $142,200.00 per year, depending on experience, location, and employer.

What is ML inference?

ML inference refers to the process of using a trained machine learning model to make predictions or decisions based on new data. After a model has been trained on historical data, inference is the phase where that model is deployed and used in real-world applications, such as recognizing speech, detecting objects in images, or recommending products. The focus in ML inference is on speed, efficiency, and scalability to ensure quick predictions, often in real time. This process is critical for practical applications like mobile apps, web services, and embedded systems. Optimizing inference involves reducing latency, memory usage, and computational requirements.

What is the difference between Ml Inference vs Data Scientist?

AspectML InferenceData Scientist
Required CredentialsKnowledge of machine learning models, programming skillsDegree in data science, statistics, or related fields
Work EnvironmentDeploying models in production, real-time data processingData analysis, model development, research
Industry UsageAI product deployment, software companiesResearch institutions, tech firms, consulting

ML Inference focuses on deploying trained models to make predictions on new data, often in real-time. Data Scientists develop and analyze models, working primarily in research and development. While both roles require understanding of machine learning, ML Inference emphasizes deployment and operationalization, whereas Data Scientists focus on model creation and analysis.

What are some common challenges faced by ML inference engineers when deploying models to production?

ML Inference Engineers often encounter challenges such as optimizing model latency and throughput to meet production requirements, ensuring compatibility with diverse hardware environments, and managing model versioning and updates without disrupting service. Additionally, balancing resource utilization and inference accuracy while monitoring real-time performance metrics is crucial. Collaboration with data scientists, DevOps, and software engineers is typically essential to streamline deployment and maintain robust, scalable inference pipelines.

What are the key skills and qualifications needed to thrive in ML inference?

To thrive in ML Inference, you need a solid background in machine learning principles, programming (Python or C++), and experience with deploying models at scale, often supported by a degree in computer science or a related field. Familiarity with frameworks and tools such as TensorFlow, PyTorch, ONNX, and cloud platforms like AWS SageMaker or Google AI Platform is typically required. Strong problem-solving skills, attention to detail, and effective communication are crucial soft skills for collaborating with multidisciplinary teams and optimizing model performance. These skills ensure efficient, scalable, and reliable deployment of machine learning solutions in real-world applications.

Is ML inference a high paying job?

ML inference roles are generally well-paying, especially for those with skills in machine learning frameworks, programming, and cloud platforms. Salaries vary based on experience, location, and industry, but they tend to be higher than average for tech-related positions.
What are popular job titles related to Ml Inference jobs in Newark, NJ? For Ml Inference jobs in Newark, NJ, the most frequently searched job titles are:
What job categories do people searching Ml Inference jobs in Newark, NJ look for? The top searched job categories for Ml Inference jobs in Newark, NJ are:
What cities near Newark, NJ are hiring for Ml Inference jobs? Cities near Newark, NJ with the most Ml Inference job openings:
Infographic showing various Ml Inference job openings in Newark, NJ as of July 2026, with employment types broken down into 96% Full Time, 2% Part Time, and 2% Contract. Highlights an 81% Physical, 5% Hybrid, and 14% Remote job distribution, with an average salary of $128,351 per year, or $61.7 per hour.

Solution Architect

The Consensus

Manhattan, NY • On-site

$120 - $150/hr

Other

Medical, Dental, Vision, Retirement, PTO

Posted 6 days ago


Job description

ABOUT BASETEN

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting‑edge models into production. We're growing quickly and recently raised our $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products.

THE ROLE

As a Solution Architect at Baseten you will partner closely with Sales and customers to translate business needs into technical solutions, run technical discovery, and guide repeatable deployments and proofs of value for customers. This role is a great fit for entrepreneurial, customer-facing technical professionals who want a front-row view into how modern companies adopt AI at scale, and who enjoy working across technical discovery, solution design, demos, deployment scoping, and hands‑on customer implementations, in close partnership with Sales and Engineering.

RESPONSIBILITIES
  • Partner with Sales on customer discovery calls (most often second calls, occasionally first calls for large accounts).

  • Lead demos and technical scoping to align on success criteria, architecture, and deployment approach.

  • Own benchmarking and repeatable deployments, including:

  • Handling standard deployment patterns and configurations across many modalities – LLMs, embeddings, image and video generation, VoiceAI, etc.

  • Advising on tradeoffs like H100s vs B200s and latency‑optimized vs throughput‑optimized setups.

  • Driving consistent “playbook” style deployments for common models and use cases.

  • Become a power user of different runtimes such as vllm, sglang, and TRT‑LMM and all the common configurations and tradeoffs between them.

  • Drive POC and project execution, including:

  • Scoping POCs and keeping stakeholders aligned on timeline, deliverables, and next steps.

  • Acting as the “ringleader” or project manager for POCs.

  • Pulling in Forward Deployed Engineering (FDE) support when deeper or more complex technical work is needed.

REQUIREMENTS
  • AI/ML background and the ability to credibly discuss AI/ML topics with technical stakeholders.

  • Strong customer‑facing communication skills, including the ability to run structured discovery and clarify ambiguous requirements.

  • Technical depth to scope solutions, without needing to write production code.

  • Ability to script and prototype as needed, including comfort “vibe coding” to move quickly in technical workflows.

NICE TO HAVE
  • Experience running or supporting benchmarks for ML inference deployments.

  • Familiarity with infrastructure tradeoffs relevant to inference performance and cost (for example GPU selection and latency versus throughput tuning).

  • Experience serving as a cross‑functional technical lead for customer POCs, including coordination across Sales and Engineering.

BENEFITS
  • Competitive compensation, including meaningful equity.

  • 100% coverage of medical, dental, and vision insurance for employee and dependents.

  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day).

  • Paid parental leave.

  • Fertility and family‑building stipend through Carrot.

  • Company‑facilitated 401(k).

  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward‑thinking team, we would love to hear from you.

At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

#J-18808-Ljbffr