1

Python Llm Jobs in Napa, CA (NOW HIRING)

Architect and implement robust, scalable solutions using Python, Langchain/LangGraph, and LLM frameworks. * Act as a trusted technical advisor to customers, understanding their needs and crafting ...

You are highly comfortable working with Python, LLM APIs (OpenAI, Anthropic, open-source), frameworks like LangChain/LlamaIndex, and modern deployment stacks (Vercel, Next.js). You have a GitHub full ...

... Python and SQL - Experience with Docker and containerized deployments - Skilled in AI techniques ... LLM optimization - Implementing data integration solutions using AWS, Azure, GCP - Utilizing AWS ...

Showing results 21-40

Python Llm information

See Napa, CA salary details

$14

$66

$97

How much do python llm jobs pay per hour?

As of Aug 23, 2026, the average hourly pay for python llm in Napa, CA is $66.45, according to ZipRecruiter salary data. Most workers in this role earn between $54.76 and $75.48 per hour, depending on experience, location, and employer.

What is a Python LLM?

A Python LLM job involves working with Large Language Models (LLMs) using Python to develop, fine-tune, and deploy AI models. Responsibilities may include data preprocessing, prompt engineering, model optimization, and integration with applications. Professionals in this role often work with frameworks like TensorFlow, PyTorch, or Hugging Face Transformers. They may also contribute to improving model efficiency, reducing bias, and ensuring ethical AI usage.

What are the key skills and qualifications needed to thrive in the Python LLM position, and why are they important?

To excel as a Python LLM (Large Language Model) Engineer, you need strong skills in Python programming, machine learning, and natural language processing, typically supported by a degree in computer science or a related field. Proficiency with libraries such as TensorFlow, PyTorch, Hugging Face Transformers, and experience with model deployment platforms are often essential, alongside certifications in AI or data science. Effective communication, problem-solving abilities, and collaboration are important soft skills for working in interdisciplinary teams and delivering results in dynamic environments. These skills ensure the development, fine-tuning, and deployment of advanced language models that meet both technical and business objectives.

What are some common challenges faced by Python LLM engineers in their daily work?

Python LLM Engineers often encounter challenges related to optimizing model performance, managing large datasets, and adapting models to specific business needs. Working with large-scale language models requires balancing computational resource limitations with the need for high accuracy and efficiency. Collaboration with data scientists, product managers, and DevOps engineers is routine to ensure seamless model integration and deployment. Staying updated on the latest advancements in NLP and continuously improving models based on user feedback are also important aspects of the role.

What cities near Napa, CA are hiring for Python Llm jobs?

Cities near Napa, CA with the most Python Llm job openings:

Infographic showing various Python Llm job openings in Napa, CA as of June 2026, with employment types broken down into 91% Full Time, 5% Part Time, and 4% Contract. Highlights an 82% Physical, 5% Hybrid, and 13% Remote job distribution, with an average salary of $138,208 per year, or $66.4 per hour.

LLM Inference Frameworks and Optimization Engineer

Together AI

San Francisco, CA • On-site

$160K - $230K/yr

Full-time

Medical

Re-posted 3 days ago


Job description

About the Role
At Together.ai, we are building state-of-the-art infrastructure to enable efficient and scalable inference for large language models (LLMs). Our mission is to optimize inference frameworks, algorithms, and infrastructure, pushing the boundaries of performance, scalability, and cost-efficiency.
We are seeking anInference Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines that support multimodal and language models at scale. This role will focus on low-latency, high-throughput inference, GPU/accelerator optimizations, and software-hardware co-design, ensuring efficient large-scale deployment of LLMs and vision models.
This role offers a unique opportunity to shape the future of LLM inference infrastructure, ensuring scalable, high-performance AI deployment across a diverse range of applications. If you're passionate about pushing the boundaries of AI inference, we'd love to hear from you!
Responsibilities
Inference Framework Development and Optimization
  • Design and develop fault-tolerant, high-concurrency distributed inference engine for text, image, and multimodal generation models.
  • Implement and optimize distributed inference strategies, including Mixture of Experts (MoE) parallelism, tensor parallelism, pipeline parallelism for high-performance serving.
  • Apply CUDA graph optimizations, TensorRT/TRT-LLM graph optimizations, and PyTorch-based compilation (torch.compile), and speculative decoding to enhance efficiency and scalability.
Software-Hardware Co-Design and AI Infrastructure
  • Collaborate with hardware teams on performance bottleneck analysis, co-optimize inference performance for GPUs, TPUs, or custom accelerators.
  • Work closely with AI researchers and infrastructure engineers to develop efficient model execution plans and optimize E2E model serving pipelines.
Requirements
Must-Have:
  • Experience:
    • 3+ years of experience in deep learning inference frameworks, distributed systems, or high-performance computing.
  • Technical Skills:
    • Familiar with at least one LLM inference frameworks (e.g., TensorRT-LLM, vLLM, SGLang, TGI(Text Generation Inference)).
    • Background knowledge and experience in at least one of the following: GPU programming (CUDA/Triton/TensorRT), compiler, model quantization, and GPU cluster scheduling.
    • Deep understanding of KV cache systems like Mooncake, PagedAttention, or custom in-house variants.
  • Programming:
    • Proficient in Python and C++/CUDA for high-performance deep learning inference.
  • Optimization Techniques:
    • Deep understanding of Transformer architectures and LLM/VLM/Diffusion model optimization.
    • Knowledge of inference optimization, such as workload scheduling, CUDA graph, compiled, efficient kernels
  • Soft Skills:
    • Strong analytical problem-solving skills with a performance-driven mindset.
    • Excellent collaboration and communication skills across teams.

Nice-to-Have:
  • Experience in developing software systems for large-scale data center networks with RDMA/RoCE
  • Familiar with distributed filesystem(e.g., 3FS, HDFS, Ceph)
  • Familiar with open source distributed scheduling/orchestration frameworks, such as Kubernetes (K8S)
  • Contributions to open-source deep learning inference projects.

About Together AI
Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers in our journey in building the next generation AI infrastructure.
Compensation
We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $160,000 - $230,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.
Equal Opportunity
Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
Please see our privacy policy at https://www.together.ai/privacy