1

Ml Inference Jobs in Portland, OR (NOW HIRING)

Solid understanding of ML fundamentals, model parallelism and inference serving techniques. Proficiency in Python (and optionally C++) for simulator design and data analysis. 3+ years of hands-on ...

Integrate and fine-tune Large Language Models (LLMs) and other AI/ML models into enterprise applications. Develop and implement strategies for model deployment, inference, and monitoring, with an ...

... ML infrastructure * MLOps practices (CI/CD, monitoring, model versioning) * Knowledge of: * Signal processing or physics-based modeling * Graph-based reasoning or causal inference * Full software ...

Showing results 21-40

Ml Inference information

See Portland, OR salary details

$39.8K

$130.2K

$208.4K

How much do ml inference jobs pay per year?

As of Sep 11, 2026, the average yearly pay for ml inference in Portland, OR is $130,165.00, according to ZipRecruiter salary data. Most workers in this role earn between $104,500.00 and $144,200.00 per year, depending on experience, location, and employer.

What is ML inference?

ML inference refers to the process of using a trained machine learning model to make predictions or decisions based on new data. After a model has been trained on historical data, inference is the phase where that model is deployed and used in real-world applications, such as recognizing speech, detecting objects in images, or recommending products. The focus in ML inference is on speed, efficiency, and scalability to ensure quick predictions, often in real time. This process is critical for practical applications like mobile apps, web services, and embedded systems. Optimizing inference involves reducing latency, memory usage, and computational requirements.

What are the key skills and qualifications needed to thrive in ML inference?

To thrive in ML Inference, you need a solid background in machine learning principles, programming (Python or C++), and experience with deploying models at scale, often supported by a degree in computer science or a related field. Familiarity with frameworks and tools such as TensorFlow, PyTorch, ONNX, and cloud platforms like AWS SageMaker or Google AI Platform is typically required. Strong problem-solving skills, attention to detail, and effective communication are crucial soft skills for collaborating with multidisciplinary teams and optimizing model performance. These skills ensure efficient, scalable, and reliable deployment of machine learning solutions in real-world applications.

What are some common challenges faced by ML inference engineers when deploying models to production?

ML Inference Engineers often encounter challenges such as optimizing model latency and throughput to meet production requirements, ensuring compatibility with diverse hardware environments, and managing model versioning and updates without disrupting service. Additionally, balancing resource utilization and inference accuracy while monitoring real-time performance metrics is crucial. Collaboration with data scientists, DevOps, and software engineers is typically essential to streamline deployment and maintain robust, scalable inference pipelines.

What is the difference between Ml Inference vs Data Scientist?

AspectML InferenceData Scientist
Required CredentialsKnowledge of machine learning models, programming skillsDegree in data science, statistics, or related fields
Work EnvironmentDeploying models in production, real-time data processingData analysis, model development, research
Industry UsageAI product deployment, software companiesResearch institutions, tech firms, consulting

ML Inference focuses on deploying trained models to make predictions on new data, often in real-time. Data Scientists develop and analyze models, working primarily in research and development. While both roles require understanding of machine learning, ML Inference emphasizes deployment and operationalization, whereas Data Scientists focus on model creation and analysis.

What are popular job titles related to Ml Inference jobs in Portland, OR?

For Ml Inference jobs in Portland, OR, the most frequently searched job titles are:

What cities near Portland, OR are hiring for Ml Inference jobs?

Cities near Portland, OR with the most Ml Inference job openings:

Infographic showing various Ml Inference job openings in Portland, OR as of September 2026, with employment types broken down into 56% Full Time, and 44% Contract. Highlights an 60% In-person, and 40% Remote job distribution, with an average salary of $130,165 per year, or $62.6 per hour.

Senior Performance Architect, Nemotron

Hillsboro, OR • On-site

Nvidia
Computer and Electronic Product Manufacturing • 10K+ employees

$181K/yr

Full-time

Re-posted 24 days ago


Nvidia rating

9.6

Company rating: 9.6 out of 10

Based on 18 frontline employees who took The Breakroom Quiz


Job description

We are now looking for a Senior Performance Architect for Nemotron. At NVIDIA, we are redefining the future of AI systems through deep model-system-hardware co-design. We are looking for a forward-thinking Nemotron Performance Architect to shape the next generation of Nemotron models through performance modeling, analysis, and forward projections.

In this role, you will predict before we build - developing high-fidelity models to evaluate how architectural choices translate into real-world deployment efficiency. You will ensure that future models achieve Pareto-optimal trade-offs across accuracy, throughput, and interactivity on target platforms. Recent efforts such as LatentMoE architectures and the Nemotron Super model exemplify the kind of performance-driven co-design you will help advance-where modeling insights directly shape model architecture and system efficiency at scale.

This role sits at the center of Generative AI evolution, partnering across research, framework development, compiler, and hardware teams to guide decisions that determine how efficiently intelligence scales in production. What You'll Be Doing: Develop high-fidelity analytical performance models to prototype emerging algorithmic techniques & hardware optimizations to drive model-hardware co-design Nemotron family of models. Prioritize features to guide future software and hardware roadmap based on detailed performance modeling and analysis Model end-to-end performance impact of emerging GenAI workflows - such as Speculative Decoding, Agentic Pipelines, Inference-time compute scaling, RL etc.

- to understand future datacenter needs This position requires you to keep up with the latest DL research and collaborate with diverse teams, including DL researchers, hardware architects, and software engineers. What we need to see: A minimum qualification of a Master's degree (or equivalent experience) in Computer Science, Electrical Engineering or related fields. Strong background in computer architecture, roofline modeling, queuing theory and statistical performance analysis techniques.

Solid understanding of ML fundamentals, model parallelism and inference serving techniques. Proficiency in Python (and optionally C++) for simulator design and data analysis. 3+ years of hands-on experience in system evaluation of AI/ML workloads or performance analysis, modeling and optimizations for AI.

Comfortable defining metrics, designing experiments and visualizing large performance datasets to identify resource bottlenecks. Experience with deep learning frameworks like PyTorch, TRT-LLM, VLLM, SGLang A Growth mindset and pragmatic "measure, iterate, deliver" approach. Ways to Stand Out from the Crowd Proven track record of working in multi-functional teams, spanning algorithms, software and hardware architecture.

Ability to distill complex analyses into clear recommendations for both technical and non-technical collaborators. Experience with GPU computing (CUDA) Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

You will also be eligible for equity and benefits. Applications for this job will be accepted at least until May 23, 2026. This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.


What Nvidia employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Nvidia logo

About Nvidia

Sourced by ZipRecruiter

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.

Industry

Computer and electronic product manufacturing

Company size

10,000+ Employees

Headquarters location

Santa Clara, CA, US