1

Ml Inference Jobs in Buffalo, NY (NOW HIRING)

Lead the deployment of ML models into scalable, reliable production systems ... Architect training, inference, and evaluation pipelines for structured and unstructured data * Own ...

Experience optimizing highโ€‘latency models for realโ€‘time inference. * Backend software engineering experience in the cloud (AWS / GCP) with a focus on microservices (docker) and the ML model ...

Machine Learning Engineer III, Data

Buffalo, NY ยท On-site

$110K - $133K/yr

Experience optimizing high-latency models for real-time inference * Backend software engineering experience in the cloud (AWS / GCP) with a focus on microservices (docker) and the ML model ...

Experience optimizing high-latency models for real-time inference * Backend software engineering experience in the cloud (AWS / GCP) with a focus on microservices (docker) and the ML model ...

Experience optimizing high-latency models for real-time inference * Backend software engineering experience in the cloud (AWS / GCP) with a focus on microservices (docker) and the ML model ...

Experience optimizing high-latency models for real-time inference * Backend software engineering experience in the cloud (AWS / GCP) with a focus on microservices (docker) and the ML model ...

Machine Learning Engineer IV, Data

Buffalo, NY ยท On-site

$54 - $71.50/hr

Experience optimizing high-latency models for real-time inference * Backend software engineering experience in the cloud (AWS / GCP) with a focus on microservices (docker) and the ML model ...

Experience optimizing high-latency models for real-time inference * Backend software engineering experience in the cloud (AWS / GCP) with a focus on microservices (docker) and the ML model ...

Senior AI Engineer

Boston, NY ยท On-site

$200K - $350K/yr

Work with LLMs, generative AI, and modern ML frameworks. * Build inference and evaluation pipelines. * Optimize model performance, latency, and cost. * Integrate AI capabilities into production ...

Staff AI Engineer

Boston, NY ยท On-site

$200K - $350K/yr

Deep Python and ML engineering expertise. * Strong production LLM experience. * Strong distributed systems and system design skills. * Experience optimizing inference systems. * Strong technical ...

Principal Founding Product Engineer

Boston, NY ยท On-site

$200K - $350K/yr

... AI/ML or software engineering experience. * Deep expertise in production AI systems. * Strong distributed systems and architecture skills. * Significant LLM and inference experience. * Strong ...

AI Founder, Edge Hardware

Boston, NY ยท On-site +1

$250K/yr

Early-stage operator experience with a proven track record of building, selling, and navigating the edge ML and on-device inference world. * You may have shipped models across multiple NPU or edge ...

AI Founder, AI Compute

Boston, NY ยท On-site +1

$250K/yr

You may have run ML platform or infrastructure teams deploying inference across cloud, neocloud, and owned GPU or edge hardware. * You may have sold into VP Platform, data center, or infrastructure ...

Machine Learning Engineer IV, Data

Buffalo, NY ยท On-site

$140K - $180K/yr

Experience optimizing high-latency models for real-time inference * Backend software engineering experience in the cloud (AWS / GCP) with a focus on microservices (docker) and the ML model ...

next page

Showing results 1-20

Ml Inference information

See Buffalo, NY salary details

$36.3K

$118.9K

$190.3K

How much do ml inference jobs pay per year?

As of Aug 21, 2026, the average yearly pay for ml inference in Buffalo, NY is $118,892.00, according to ZipRecruiter salary data. Most workers in this role earn between $95,400.00 and $131,700.00 per year, depending on experience, location, and employer.

What is ML inference?

ML inference refers to the process of using a trained machine learning model to make predictions or decisions based on new data. After a model has been trained on historical data, inference is the phase where that model is deployed and used in real-world applications, such as recognizing speech, detecting objects in images, or recommending products. The focus in ML inference is on speed, efficiency, and scalability to ensure quick predictions, often in real time. This process is critical for practical applications like mobile apps, web services, and embedded systems. Optimizing inference involves reducing latency, memory usage, and computational requirements.

What are the key skills and qualifications needed to thrive in ML inference?

To thrive in ML Inference, you need a solid background in machine learning principles, programming (Python or C++), and experience with deploying models at scale, often supported by a degree in computer science or a related field. Familiarity with frameworks and tools such as TensorFlow, PyTorch, ONNX, and cloud platforms like AWS SageMaker or Google AI Platform is typically required. Strong problem-solving skills, attention to detail, and effective communication are crucial soft skills for collaborating with multidisciplinary teams and optimizing model performance. These skills ensure efficient, scalable, and reliable deployment of machine learning solutions in real-world applications.

What are some common challenges faced by ML inference engineers when deploying models to production?

ML Inference Engineers often encounter challenges such as optimizing model latency and throughput to meet production requirements, ensuring compatibility with diverse hardware environments, and managing model versioning and updates without disrupting service. Additionally, balancing resource utilization and inference accuracy while monitoring real-time performance metrics is crucial. Collaboration with data scientists, DevOps, and software engineers is typically essential to streamline deployment and maintain robust, scalable inference pipelines.

What is the difference between Ml Inference vs Data Scientist?

AspectML InferenceData Scientist
Required CredentialsKnowledge of machine learning models, programming skillsDegree in data science, statistics, or related fields
Work EnvironmentDeploying models in production, real-time data processingData analysis, model development, research
Industry UsageAI product deployment, software companiesResearch institutions, tech firms, consulting

ML Inference focuses on deploying trained models to make predictions on new data, often in real-time. Data Scientists develop and analyze models, working primarily in research and development. While both roles require understanding of machine learning, ML Inference emphasizes deployment and operationalization, whereas Data Scientists focus on model creation and analysis.

What are popular job titles related to Ml Inference jobs in Buffalo, NY?

For Ml Inference jobs in Buffalo, NY, the most frequently searched job titles are:

What job categories do people searching Ml Inference jobs in Buffalo, NY look for?

The top searched job categories for Ml Inference jobs in Buffalo, NY are:

What cities near Buffalo, NY are hiring for Ml Inference jobs?

Cities near Buffalo, NY with the most Ml Inference job openings:

Infographic showing various Ml Inference job openings in Buffalo, NY as of August 2026, with employment types broken down into 89% Full Time, 7% Part Time, and 4% Contract. Highlights an 79% Physical, 6% Hybrid, and 15% Remote job distribution, with an average salary of $118,892 per year, or $57.2 per hour.

Principal AI/ML Engineer

Pragmatike

Boston, NY โ€ข On-site

Full-time

Medical, Dental, Vision, Retirement

Re-posted 24 days ago


Job description

Location: Cambridge, MA (Eastern Time / UTC -4) Relocation package available
Relocation package: Available
Start date: ASAP
Languages: English (required)
About the Role
Pragmatike is hiring on behalf of a fast-growing AI startup recognized as a Top 10 GenAI company by GTM Capital, founded by MIT CSAIL researchers.
We are looking for a Senior AI/ML Engineer to join a high-performing AI engineering team building and deploying real-world AI systems used by Fortune 500 customers. You will take technical ownership of critical ML components, lead model development and productionization efforts, and work closely with AI researchers, platform engineers, and product teams to deliver scalable, production-grade AI systems.
This role is intended for experienced AI/ML engineers who have already built and deployed models in production environments and want to operate at the intersection of research-quality AI and real-world systems engineering.
What You'll Do
  • Design, build, train, and optimize machine learning and deep learning models for production use
  • Lead the deployment of ML models into scalable, reliable production systems
  • Architect training, inference, and evaluation pipelines for structured and unstructured data
  • Own model performance, reliability, scalability, and lifecycle management
  • Drive experimentation, model iteration, and performance benchmarking
  • Implement and maintain high-quality ML codebases primarily in Python
  • Monitor, debug, and improve live production models
  • Collaborate with platform, backend, and infrastructure teams on system integration
  • Contribute to technical direction, architecture decisions, and AI roadmap planning
  • Mentor junior engineers and support team-level technical excellence
What We're Looking For
  • Bachelor's, Master's, or PhD in Computer Science, Engineering, Mathematics, or a related field
  • 7+ years of professional experience in AI/ML, data science, or applied machine learning roles
  • Proven experience deploying ML systems into production environments
  • Strong understanding of ML fundamentals, evaluation methodologies, and model lifecycle management
  • Deep hands-on experience with ML frameworks such as PyTorch, TensorFlow, or scikit-learn
  • Strong proficiency in Python
  • Experience designing scalable ML systems and pipelines
  • Ability to operate independently and take technical ownership of complex systems
  • Strong problem-solving, architecture, and system-design skills
Bonus Points
  • Experience with deep learning, LLMs, GenAI, or modern AI architectures
  • Experience building AI systems used by enterprise or large-scale customers
  • Strong experience with cloud platforms (AWS, GCP, Azure)
  • MLOps experience (CI/CD for ML, model deployment, monitoring, observability)
  • Research background, publications, or applied AI innovation projects
  • Experience leading AI initiatives or mentoring ML engineers
Why This Role Is High-Impact
  • Work directly with MIT CSAIL-founded leadership
  • Build AI systems deployed to Fortune 500 customers
  • Operate at the intersection of cutting-edge AI and real-world production systems
  • High ownership and influence over technical direction
  • Join a well-funded, fast-growing AI startup with strong technical ambition
  • Clear path toward Staff/Principal AI Engineer or AI Architect roles
Benefits
  • Competitive salary & equity options
  • Sign-on bonus
  • Health, Dental, and Vision
  • 401(k)

Pragmatike is an Equal Opportunity Employer and is committed to providing equal employment opportunities to all applicants without discrimination. We recruit on behalf of our clients and prohibit discrimination and harassment based on race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state, or local laws. This policy applies to all terms and conditions of employment, including recruiting, hiring, placement, promotion, termination, layoff, recall, transfer, leaves of absence, compensation, and training.We are committed to a fair and inclusive hiring process. We process your personal data solely for recruitment purposes, in accordance with applicable privacy laws, and maintain reasonable safeguards to protect your information. Your data may be shared with our client(s) for hiring consideration, but will not be disclosed to third parties outside of the recruitment process.