1

Ml Inference Jobs in Portland, OR (NOW HIRING)

Integrate and fine-tune Large Language Models (LLMs) and other AI/ML models into enterprise applications. Develop and implement strategies for model deployment, inference, and monitoring, with an ...

AI Red Team Lead Engineer

Gresham, OR · On-site

$108K - $143K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

... on AI/ML systems, platforms, and integrations, in addition to traditional enterprise attack ... Training, evaluation, and inference pipelines * Data ingestion, labeling, and governance controls

... ML infrastructure * MLOps practices (CI/CD, monitoring, model versioning) * Knowledge of: * Signal processing or physics-based modeling * Graph-based reasoning or causal inference * Full software ...

Senior Manager - Advanced Technology

Portland, OR · On-site

  • Medical

  • Dental

  • Vision

  • Retirement

Architect and deliver compute stacks for AI/ML (training/inference), physicsbased modeling, and streaming analytics, prioritizing reliability, scalability, and total cost of ownership. * Lead a team ...

Senior Manager - Advanced Technology

Portland, OR · On-site

  • Medical

  • Dental

  • Vision

  • Retirement

Architect and deliver compute stacks for AI/ML (training/inference), physics-based modeling, and streaming analytics, prioritizing reliability, scalability, and total cost of ownership. * Lead a team ...

Showing results 21-38

Ml Inference information

See Portland, OR salary details

$39.8K

$130.2K

$208.4K

How much do ml inference jobs pay per year?

As of Aug 13, 2026, the average yearly pay for ml inference in Portland, OR is $130,165.00, according to ZipRecruiter salary data. Most workers in this role earn between $104,500.00 and $144,200.00 per year, depending on experience, location, and employer.

What is ML inference?

ML inference refers to the process of using a trained machine learning model to make predictions or decisions based on new data. After a model has been trained on historical data, inference is the phase where that model is deployed and used in real-world applications, such as recognizing speech, detecting objects in images, or recommending products. The focus in ML inference is on speed, efficiency, and scalability to ensure quick predictions, often in real time. This process is critical for practical applications like mobile apps, web services, and embedded systems. Optimizing inference involves reducing latency, memory usage, and computational requirements.

What is the difference between Ml Inference vs Data Scientist?

AspectML InferenceData Scientist
Required CredentialsKnowledge of machine learning models, programming skillsDegree in data science, statistics, or related fields
Work EnvironmentDeploying models in production, real-time data processingData analysis, model development, research
Industry UsageAI product deployment, software companiesResearch institutions, tech firms, consulting

ML Inference focuses on deploying trained models to make predictions on new data, often in real-time. Data Scientists develop and analyze models, working primarily in research and development. While both roles require understanding of machine learning, ML Inference emphasizes deployment and operationalization, whereas Data Scientists focus on model creation and analysis.

What are some common challenges faced by ML inference engineers when deploying models to production?

ML Inference Engineers often encounter challenges such as optimizing model latency and throughput to meet production requirements, ensuring compatibility with diverse hardware environments, and managing model versioning and updates without disrupting service. Additionally, balancing resource utilization and inference accuracy while monitoring real-time performance metrics is crucial. Collaboration with data scientists, DevOps, and software engineers is typically essential to streamline deployment and maintain robust, scalable inference pipelines.

What are the key skills and qualifications needed to thrive in ML inference?

To thrive in ML Inference, you need a solid background in machine learning principles, programming (Python or C++), and experience with deploying models at scale, often supported by a degree in computer science or a related field. Familiarity with frameworks and tools such as TensorFlow, PyTorch, ONNX, and cloud platforms like AWS SageMaker or Google AI Platform is typically required. Strong problem-solving skills, attention to detail, and effective communication are crucial soft skills for collaborating with multidisciplinary teams and optimizing model performance. These skills ensure efficient, scalable, and reliable deployment of machine learning solutions in real-world applications.

Is ML inference a high paying job?

ML inference roles are generally well-paying, especially for those with skills in machine learning frameworks, programming, and cloud platforms. Salaries vary based on experience, location, and industry, but they tend to be higher than average for tech-related positions.
What are popular job titles related to Ml Inference jobs in Portland, OR? For Ml Inference jobs in Portland, OR, the most frequently searched job titles are:
What job categories do people searching Ml Inference jobs in Portland, OR look for? The top searched job categories for Ml Inference jobs in Portland, OR are:
What cities near Portland, OR are hiring for Ml Inference jobs? Cities near Portland, OR with the most Ml Inference job openings:
Infographic showing various Ml Inference job openings in Portland, OR as of August 2026, with employment types broken down into 91% Full Time, 5% Part Time, and 4% Contract. Highlights an 79% Physical, 6% Hybrid, and 15% Remote job distribution, with an average salary of $130,165 per year, or $62.6 per hour.

Google AI Architect

Deloitte

Portland, OR • On-site

Full-time

Posted 6 days ago


Deloitte rating

8.2

Company rating: 8.2 out of 10

Based on 92 frontline employees who took The Breakroom Quiz

45th of 150 rated financial services


Job description

Google AI Architect/AI and Engineering

Join our AI & Engineering team in transforming technology platforms, driving innovation, and helping make a significant impact on our clients' success. You'll work alongside talented professionals reimagining and re-engineering operations and processes that are critical to businesses. Your contributions can help clients improve financial performance, accelerate new digital ventures, and fuel growth through innovation.
AI & Engineering leverages cutting-edge engineering capabilities to build, deploy, and operate integrated/verticalized sector solutions in software, data, AI, network, and hybrid cloud infrastructure. These solutions are powered by engineering for business advantage, transforming mission-critical operations. We enable clients to stay ahead with the latest advancements by transforming engineering teams and modernizing technology & data platforms. Our delivery models are tailored to meet each client's unique requirements.
Engineering as a Service provides complete design, implementation, and technology operations, leveraging our core engineering expertise. We transform engineering teams, modernize technology, and deliver complex programs with a product engineering approach. Our flexible delivery models-traditional teams, pools, or pods-are tailored to each client's needs, offering engineering-led advisory, implementation, and operational capabilities to accelerate innovation.

Recruiting for this role ends on 10-31-2026
Work you'll do:

  • Architect and deliver enterprise AI platforms and applications on Google Cloud using Vertex AI and Gemini; optimize for scalability, reliability, security, and cost.
  • Design, fine-tune, evaluate, and govern LLM solutions with Gemini on Vertex AI (prompt/tool/function calling, safety policies, Vector Search, evaluation); implement deployment, inference optimization, and monitoring.
  • Build RAG and agentic solutions using Vertex AI Vector Search and BigQuery vector; implement context management, retrieval strategies, and observability.
  • Define end-to-end architectures across data pipelines, feature engineering, model lifecycle, APIs/microservices, and CI/CD/MLOps/LLMOps with Vertex AI Pipelines and Cloud Build.
  • Lead cloud-native development on GKE, Cloud Run, Pub/Sub, BigQuery, Cloud SQL/Spanner, Memorystore, and Terraform; enforce application and agentic design patterns.
  • Implement security and governance for AI/ML systems (data privacy, model poisoning, adversarial attacks); apply Gemini safety features and enterprise guardrails.

Responsibilities include:

  • Architect and Design: Design and development of enterprise-grade AI applications and platforms, with a focus on scaling AI solutions for production. This includes defining the technical architecture, selecting appropriate technologies, and ensuring solutions are robust, scalable, and secure.
  • LLM and AI Integration: Integrate and fine-tune Large Language Models (LLMs) and other AI/ML models into enterprise applications. Develop and implement strategies for model deployment, inference, and monitoring, with an emphasis on production-level performance and reliability.
  • Enterprise Architecture: Collaborate with enterprise architects to ensure AI solutions align with the broader company's technical strategy, governance, and standards.
  • Cloud and GenAI Native Development: Design and deploy applications using Cloud Native principles on a hyperscaler platform (AWS, Azure, GCP). Leverage a wide range of hyperscaler tools and services, including containers (Docker, Kubernetes), serverless functions, and managed databases. Should have experience in leveraging various GenAI tools to accelerate software development life cycle.
  • Security & Governance: Ensure the security of all AI/ML systems by addressing potential vulnerabilities such as data privacy concerns, model poisoning, and adversarial attacks.
  • Design Patterns: Apply and enforce Application Design Patterns and Agentic Design Patterns to build resilient and maintainable software systems.

 Required Qualifications

  • Bachelor's degree in Computer Science, Engineering or a related technical field.
  • 6+ years' experience as a Software or Solution Architect, with a strong focus on application development and scaling solutions for production environments.
  • 5+ years hands-on with Google Cloud, including 2+ end-to-end enterprise implementations in production.
  • 4+ years designing and implementing Google Cloud networks, security controls, and landing zones using Terraform.
  • 2+ years building and operating containerized workloads on GKE (autoscaling, ingress, monitoring/observability).
  • 2+ years implementing CI/CD and DevSecOps with Cloud Build, GitHub Actions, or Jenkins.
  • 3+ years executing migration or modernization programs to Google Cloud (rehost, replatform, refactor).
  • 2+ years applying AI/GenAI on Google Cloud with Vertex AI and Gemini, including 1+ years' production deployment (e.g. RAG with Vertex AI Search/Vector Search, prompt design, safety policies, observability).
  • Deep understanding of AI/ML concepts, including experience with LLMs and their application in enterprise settings.
  • Experience implementing multiple AI solutions in a professional, real-world environment.
  • Strong understanding of security implications related to AI/ML systems (e.g., data privacy, model poisoning, adversarial attacks).
  • Familiarity with various hyperscaler tools and services.
  • Hyperscaler Architect certification is required (e.g., AWS Certified Solutions Architect, Azure Solutions Architect Expert, or GCP Professional Cloud Architect).
  • Ability to travel up to 50% based on the work you do and the clients and industries/sectors you serve.
  • Limited immigration sponsorship may be available.

Preferred Qualifications:

  • Google Professional Machine Learning Engineer certification or the equivalent ML certification.
  • Master's degree in technology-related discipline.
  •  2+ years's leading high performance, results driven engineering teams delivering AI platforms or applications.
  • 1+ year implementing LLMOps/MLOps using Vertex AI Pipelines and Cloud Build (or similar)

Wages + Salary

The wage range for this role takes into account the wide range of factors that are considered in making compensation decisions including but not limited to skill sets; experience and training; licensure and certifications; and other business and organizational needs. The disclosed range estimate has not been adjusted for the applicable geographic differential associated with the location at which the position may be filled. At Deloitte, it is not typical for an individual to be hired at or near the top of the range for their role and compensation decisions are dependent on the facts and circumstances of each case. A reasonable estimate of the current range is $122,000-$240,500.

You may also be eligible to participate in a discretionary annual incentive program, subject to the rules governing the program, whereby an award, if any, depends on various factors, including, without limitation, individual and organizational performance.

Information for applicants with a need for accommodation: 

https://www2.deloitte.com/us/en/pages/careers/articles/join-deloitte-assistance-for-disabled-applicants.html

Qualifications:

Google AI Architect/AI and Engineering

Join our AI & Engineering team in transforming technology platforms, driving innovation, and helping make a significant impact on our clients' success. You'll work alongside talented professionals reimagining and re-engineering operations and processes that are critical to businesses. Your contributions can help clients improve financial performance, accelerate new digital ventures, and fuel growth through innovation.
AI & Engineering leverages cutting-edge engineering capabilities to build, deploy, and operate integrated/verticalized sector solutions in software, data, AI, network, and hybrid cloud infrastructure. These solutions are powered by engineering for business advantage, transforming mission-critical operations. We enable clients to stay ahead with the latest advancements by transforming engineering teams and modernizing technology & data platforms. Our delivery models are tailored to meet each client's unique requirements.
Engineering as a Service provides complete design, implementation, and technology operations, leveraging our core engineering expertise. We transform engineering teams, modernize technology, and deliver complex programs with a product engineering approach. Our flexible delivery models-traditional teams, pools, or pods-are tailored to each client's needs, offering engineering-led advisory, implementation, and operational capabilities to accelerate innovation.

Recruiting for this role ends on 10-31-2026
Work you'll do:

  • Architect and deliver enterprise AI platforms and applications on Google Cloud using Vertex AI and Gemini; optimize for scalability, reliability, security, and cost.
  • Design, fine-tune, evaluate, and govern LLM solutions with Gemini on Vertex AI (prompt/tool/function calling, safety policies, Vector Search, evaluation); implement deployment, inference optimization, and monitoring.
  • Build RAG and agentic solutions using Vertex AI Vector Search and BigQuery vector; implement context management, retrieval strategies, and observability.
  • Define end-to-end architectures across data pipelines, feature engineering, model lifecycle, APIs/microservices, and CI/CD/MLOps/LLMOps with Vertex AI Pipelines and Cloud Build.
  • Lead cloud-native development on GKE, Cloud Run, Pub/Sub, BigQuery, Cloud SQL/Spanner, Memorystore, and Terraform; enforce application and agentic design patterns.
  • Implement security and governance for AI/ML systems (data privacy, model poisoning, adversarial attacks); apply Gemini safety features and enterprise guardrails.

Responsibilities include:

  • Architect and Design: Design and development of enterprise-grade AI applications and platforms, with a focus on scaling AI solutions for production. This includes defining the technical architecture, selecting appropriate technologies, and ensuring solutions are robust, scalable, and secure.
  • LLM and AI Integration: Integrate and fine-tune Large Language Models (LLMs) and other AI/ML models into enterprise applications. Develop and implement strategies for model deployment, inference, and monitoring, with an emphasis on production-level performance and reliability.
  • Enterprise Architecture: Collaborate with enterprise architects to ensure AI solutions align with the broader company's technical strategy, governance, and standards.
  • Cloud and GenAI Native Development: Design and deploy applications using Cloud Native principles on a hyperscaler platform (AWS, Azure, GCP). Leverage a wide range of hyperscaler tools and services, including containers (Docker, Kubernetes), serverless functions, and managed databases. Should have experience in leveraging various GenAI tools to accelerate software development life cycle.
  • Security & Governance: Ensure the security of all AI/ML systems by addressing potential vulnerabilities such as data privacy concerns, model poisoning, and adversarial attacks.
  • Design Patterns: Apply and enforce Application Design Patterns and Agentic Design Patterns to build resilient and maintainable software systems.

 Required Qualifications

  • Bachelor's degree in Computer Science, Engineering or a related technical field.
  • 6+ years' experience as a Software or Solution Architect, with a strong focus on application development and scaling solutions for production environments.
  • 5+ years hands-on with Google Cloud, including 2+ end-to-end enterprise implementations in production.
  • 4+ years designing and implementing Google Cloud networks, security controls, and landing zones using Terraform.
  • 2+ years building and operating containerized workloads on GKE (autoscaling, ingress, monitoring/observability).
  • 2+ years implementing CI/CD and DevSecOps with Cloud Build, GitHub Actions, or Jenkins.
  • 3+ years executing migration or modernization programs to Google Cloud (rehost, replatform, refactor).
  • 2+ years applying AI/GenAI on Google Cloud with Vertex AI and Gemini, including 1+ years' production deployment (e.g. RAG with Vertex AI Search/Vector Search, prompt design, safety policies, observability).
  • Deep understanding of AI/ML concepts, including experience with LLMs and their application in enterprise settings.
  • Experience implementing multiple AI solutions in a professional, real-world environment.
  • Strong understanding of security implications related to AI/ML systems (e.g., data privacy, model poisoning, adversarial attacks).
  • Familiarity with various hyperscaler tools and services.
  • Hyperscaler Architect certification is required (e.g., AWS Certified Solutions Architect, Azure Solutions Architect Expert, or GCP Professional Cloud Architect).
  • Ability to travel up to 50% based on the work you do and the clients and industries/sectors you serve.
  • Limited immigration sponsorship may be available.

Preferred Qualifications:

  • Google Professional Machine Learning Engineer certification or the equivalent ML certification.
  • Master's degree in technology-related discipline.
  •  2+ years's leading high performance, results driven engineering teams delivering AI platforms or applications.
  • 1+ year implementing LLMOps/MLOps using Vertex AI Pipelines and Cloud Build (or similar)

Wages + Salary

The wage range for this role takes into account the wide range of factors that are considered in making compensation decisions including but not limited to skill sets; experience and training; licensure and certifications; and other business and organizational needs. The disclosed range estimate has not been adjusted for the applicable geographic differential associated with the location at which the position may be filled. At Deloitte, it is not typical for an ind...


What Deloitte employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom