1

Python Llm Jobs in California (NOW HIRING)

$104K - $137K/yr

Develop AI platform services and APIs using Python, FastAPI, Microservices, and Kubernetes ... LLM: vLLM, TensorRT-LLM, Triton Inference Server, Ray Serve * AI/GenAI: RAG, LangChain, LangGraph ...

$130K - $170K/yr

Develop AI platform services and APIs using Python, FastAPI, Microservices, and Kubernetes ... LLM: vLLM, TensorRT-LLM, Triton Inference Server, Ray Serve * AI/GenAI: RAG, LangChain, LangGraph ...

The ideal candidate will have strong experience working with Python, LLM solution patterns and tools (RAG, Vector DB, Agentic workflows, LoRA, etc.) cloud platforms (AWS, Databricks and Azure), and ...

Agentic AI/Python Engineer Location: Concord, CA (3days in hybrid) Duration: 12+ Months Contract ... Develop LLM-powered applications using RAG, vector search, embeddings, and related technologies.

Showing results 41-60

Python Llm information

What is a Python LLM?

A Python LLM job involves working with Large Language Models (LLMs) using Python to develop, fine-tune, and deploy AI models. Responsibilities may include data preprocessing, prompt engineering, model optimization, and integration with applications. Professionals in this role often work with frameworks like TensorFlow, PyTorch, or Hugging Face Transformers. They may also contribute to improving model efficiency, reducing bias, and ensuring ethical AI usage.

What are the key skills and qualifications needed to thrive in the Python LLM position, and why are they important?

To excel as a Python LLM (Large Language Model) Engineer, you need strong skills in Python programming, machine learning, and natural language processing, typically supported by a degree in computer science or a related field. Proficiency with libraries such as TensorFlow, PyTorch, Hugging Face Transformers, and experience with model deployment platforms are often essential, alongside certifications in AI or data science. Effective communication, problem-solving abilities, and collaboration are important soft skills for working in interdisciplinary teams and delivering results in dynamic environments. These skills ensure the development, fine-tuning, and deployment of advanced language models that meet both technical and business objectives.

What are some common challenges faced by Python LLM engineers in their daily work?

Python LLM Engineers often encounter challenges related to optimizing model performance, managing large datasets, and adapting models to specific business needs. Working with large-scale language models requires balancing computational resource limitations with the need for high accuracy and efficiency. Collaboration with data scientists, product managers, and DevOps engineers is routine to ensure seamless model integration and deployment. Staying updated on the latest advancements in NLP and continuously improving models based on user feedback are also important aspects of the role.

What are the most commonly searched types of Python Llm jobs in California?

The most popular types of Python Llm jobs in California are:

What job categories do people searching Python Llm jobs in California look for?

The top searched job categories for Python Llm jobs in California are:

What cities in California are hiring for Python Llm jobs?

Cities in California with the most Python Llm job openings:

Infographic showing various Python Llm job openings in California as of August 2026, with employment types broken down into 1% Internship, 89% Full Time, 5% Part Time, and 5% Contract. Highlights an 76% Physical, 7% Hybrid, and 17% Remote job distribution.

GenAI Engineer - LLM Infrastructure & Inference Services

West Menlo Park, CA • On-site

2T Consulting
IT Services • 51 - 200 employees

$126K - $165K/yr

Full-time

Posted 13 days ago


Job description

We are looking for a GenAI Engineer with strong expertise in LLM infrastructure, model deployment, and high-performance inference services. The ideal candidate will build and manage scalable enterprise GenAI platforms across GPU infrastructure and cloud environments.

Key Responsibilities
  • Deploy, host, and manage Large Language Models (LLMs) on GPU infrastructure for production environments.
  • Build scalable, high-performance inference services using vLLM, TensorRT-LLM, Triton Inference Server, and Ray Serve.
  • Optimize model serving for latency, throughput, GPU utilization, and cost efficiency.
  • Develop AI platform services and APIs using Python, FastAPI, Microservices, and Kubernetes.
  • Implement RAG pipelines, vector databases, and agentic AI frameworks such as LangChain and LangGraph.
  • Manage GPU infrastructure, containerization, and cloud deployments across AWS, Azure, or GCP.
  • Establish MLOps/LLMOps practices including CI/CD, model deployment, monitoring, observability, and governance.
  • Perform performance tuning, benchmarking, capacity planning, and production support for enterprise GenAI platforms.
  • Collaborate with architects, data scientists, and product teams to deliver scalable, secure, and reliable AI solutions.
Core Technologies
  • LLM: vLLM, TensorRT-LLM, Triton Inference Server, Ray Serve
  • AI/GenAI: RAG, LangChain, LangGraph, Vector Databases
  • Development: Python, FastAPI, Microservices
  • Infrastructure: Kubernetes, Docker, GPU Infrastructure
  • Cloud: AWS, Azure, GCP
  • MLOps/LLMOps: CI/CD, Monitoring, Observability, Model Deployment, Governance