1

Ml Inference Jobs in Kent, WA (NOW HIRING)

Staff Inference Engineer

Bellevue, WA · On-site

$200 - $250/hr

Inference Engineer Location: Hybrid | Bellevue, WA Area Titles: Senior and Staff (multiple roles ... Experience working at a hyperscaler, AI lab, GPU cloud provider, or large-scale ML infrastructure ...

Inference Engineer

Bellevue, WA · Remote

$117K - $140K/yr

Inference Engineer Location: Hybrid | Bellevue, WA Area Titles: Senior and Staff (multiple roles ... Experience working at a hyperscaler, AI lab, GPU cloud provider, or large-scale ML infrastructure ...

Senior Software Engineer, Inference

Bellevue, WA · On-site

$137K - $181K/yr

The Senior Software Engineer, Inference will lead designs, improve engineering standards, and ... to-end ML system performance by developing and tuning CUDA kernels, reducing model latency ...

Showing results 41-60

Ml Inference information

See Kent, WA salary details

$42.3K

$138.6K

$221.8K

How much do ml inference jobs pay per year?

As of Sep 9, 2026, the average yearly pay for ml inference in Kent, WA is $138,558.00, according to ZipRecruiter salary data. Most workers in this role earn between $111,200.00 and $153,500.00 per year, depending on experience, location, and employer.

What is ML inference?

ML inference refers to the process of using a trained machine learning model to make predictions or decisions based on new data. After a model has been trained on historical data, inference is the phase where that model is deployed and used in real-world applications, such as recognizing speech, detecting objects in images, or recommending products. The focus in ML inference is on speed, efficiency, and scalability to ensure quick predictions, often in real time. This process is critical for practical applications like mobile apps, web services, and embedded systems. Optimizing inference involves reducing latency, memory usage, and computational requirements.

What are the key skills and qualifications needed to thrive in ML inference?

To thrive in ML Inference, you need a solid background in machine learning principles, programming (Python or C++), and experience with deploying models at scale, often supported by a degree in computer science or a related field. Familiarity with frameworks and tools such as TensorFlow, PyTorch, ONNX, and cloud platforms like AWS SageMaker or Google AI Platform is typically required. Strong problem-solving skills, attention to detail, and effective communication are crucial soft skills for collaborating with multidisciplinary teams and optimizing model performance. These skills ensure efficient, scalable, and reliable deployment of machine learning solutions in real-world applications.

What are some common challenges faced by ML inference engineers when deploying models to production?

ML Inference Engineers often encounter challenges such as optimizing model latency and throughput to meet production requirements, ensuring compatibility with diverse hardware environments, and managing model versioning and updates without disrupting service. Additionally, balancing resource utilization and inference accuracy while monitoring real-time performance metrics is crucial. Collaboration with data scientists, DevOps, and software engineers is typically essential to streamline deployment and maintain robust, scalable inference pipelines.

What is the difference between Ml Inference vs Data Scientist?

AspectML InferenceData Scientist
Required CredentialsKnowledge of machine learning models, programming skillsDegree in data science, statistics, or related fields
Work EnvironmentDeploying models in production, real-time data processingData analysis, model development, research
Industry UsageAI product deployment, software companiesResearch institutions, tech firms, consulting

ML Inference focuses on deploying trained models to make predictions on new data, often in real-time. Data Scientists develop and analyze models, working primarily in research and development. While both roles require understanding of machine learning, ML Inference emphasizes deployment and operationalization, whereas Data Scientists focus on model creation and analysis.

What cities near Kent, WA are hiring for Ml Inference jobs?

Cities near Kent, WA with the most Ml Inference job openings:

Senior Software Development Engineer, AWS Mantle

Seattle, WA • On-site

Socket.dev
Network Security • 1 - 10 employees

$150 - $200/hr

Other

Medical, Dental, Vision, Life, Retirement, PTO

Posted 21 days ago


Key responsibilities

  • Design, build, and operate high-performance distributed systems that serve ML inference at massive scale across all AWS regions.

  • Own the end-to-end delivery of complex features, including requirements gathering, design, implementation, testing, deployment, and production operations.

  • Collaborate with cross-functional teams to solve problems related to capacity management, model serving, and API compatibility.


Job description

Description

We're looking for an experienced Senior Software Development Engineer to help build and scale the distributed inference engine that powers Amazon Bedrock. As part of the AWS Mantle team, you will design and deliver critical systems that enable millions of customers to access the world’s leading foundation models—securely, reliably, and at global scale. This is an opportunity to work on one of the most impactful AI infrastructure platforms at AWS, where your code and design decisions will directly shape how generative AI is served to enterprises worldwide.

Design, build, and operate high-performance distributed systems that serve ML inference at massive scale across all AWS regions.

Own the end-to-end delivery of complex features—from requirements through design, implementation, testing, deployment, and production operations.

Collaborate with cross-functional teams to solve challenging problems in capacity management, model serving, and API compatibility.

Contribute to a culture of engineering excellence by writing clean, maintainable code and driving continuous improvement in system reliability.

Influence technical direction within your team while contributing to broader architectural discussions across Mantle and Amazon Bedrock.

Key job responsibilities

As a Senior SDE on the Mantle team, you will be a hands-on technical leader who owns significant components of our inference platform. You will balance deep technical execution with thoughtful design, delivering solutions that are scalable, secure, and operationally excellent—while mentoring teammates and raising the bar for the team.

Design and implement core components of Mantle's distributed inference engine, including request routing, load balancing, model lifecycle management, and quality-of-service enforcement.

Build and operate services that onboard new foundation models rapidly while maintaining strict performance SLAs and Zero Operator Access (ZOA) security guarantees.

Drive operational excellence by owning your team's services in production—monitoring, alarming, incident response, and continuous reliability improvement.

Partner with applied scientists, ML engineers, and partner teams to integrate new model architectures and optimize inference performance across GPU/accelerator fleet.

Mentor junior and mid-level engineers through code reviews, design reviews, and hands‑on guidance that elevates the team’s technical capabilities.

About the team

The AWS Mantle team is building the next-generation inference engine that powers Amazon Bedrock—providing secure, enterprise-grade access to high-performing foundation models from the world’s leading AI companies. Our mission is to simplify and accelerate how models are served at global scale, with an unwavering commitment to customer trust through innovations like our Zero Operator Access architecture, designed so that no person—whether from AWS, a customer, or a model provider—can ever access customer inference data.

We operate at massive scale, serving inference requests across all major AWS regions with sophisticated automated capacity management and unified resource pools.

Our team values builders who thrive in ambiguity, think long-term, and are excited to define the future of AI infrastructure from the ground up.

We foster a collaborative, inclusive environment where diverse perspectives drive better solutions—and where the best ideas win regardless of where they originate.

We ship fast and iterate with purpose, having rapidly expanded from launch to supporting models from OpenAI, DeepSeek, Google, Mistral, NVIDIA, and more.

We believe work should be meaningful and fun—you’ll join a team that takes pride in making history at the forefront of generative AI.

Basic Qualifications
  • 5+ years of non‑internship professional software development experience
  • 5+ years of programming with at least one software programming language experience
  • 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • Experience as a mentor, tech lead or leading an engineering team
  • Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent experience
Preferred Qualifications
  • Master’s degree in computer science, machine learning, engineering, or related fields
  • Experience in machine learning, data mining, information retrieval, statistics or natural language processing, or experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware
  • Experience designing APIs at scale, particularly RESTful or streaming APIs with strict latency and availability requirements

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign‑on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, WA, Seattle - 168,100.00 - 227,400.00 USD annually

#J-18808-Ljbffr