1

Llm Developer Jobs in California (NOW HIRING)

LLM Agent Systems : Design and implement intelligent agent architectures for complex enterprise ... Mentor and collaborate with LLM engineers on implementation and deployment Requirements ...

Contribute to LLM engineering work that brings Tolan's AI capabilities to life. * Work cross-functionally with our frontend, design, and applied AI teams. Who We're Looking For * Relevant experience.

Data Engineer (Starlink)

Hawthorne, CA · On-site

$145K - $175K/yr

Experience with SQL and modern data tooling; exposure to Grok, internal AI platforms, or comparable LLM developer tooling is a plus * Experience with agent evaluation, guardrails, permissions, and ...

Data Engineer (Starlink)

Hawthorne, CA · On-site

$145K - $175K/yr

Experience with SQL and modern data tooling; exposure to Grok, internal AI platforms, or comparable LLM developer tooling is a plus * Experience with agent evaluation, guardrails, permissions, and ...

Experience with SQL and modern data tooling; exposure to Grok, internal AI platforms, or comparable LLM developer tooling is a plus * Experience with agent evaluation, guardrails, permissions, and ...

AI Engineer

San Francisco, CA · On-site

$200K - $250K/yr

Thinking about the developer experience - to support millions of users who leverage your work via LiteLLM's Python SDK * Building across LLM's, MCP's, and Agents by maintaining an excellent LLM ...

Developer Relations Lead

San Francisco, CA · On-site

$69.50 - $91/hr

What you will do You'll build Sapiom's presence among AI agent builders and LLM developers - serving as the connective tissue between developers and our product team. The goal: a thriving cohort of ...

Showing results 21-40

Llm Developer information

See California salary details

$25

$49

$78

How much do llm developer jobs pay per hour?

As of Aug 11, 2026, the average hourly pay for llm developer in California is $49.51, according to ZipRecruiter salary data. Most workers in this role earn between $38.89 and $60.00 per hour, depending on experience, location, and employer.

What does an LLM Developer do?

An LLM Developer designs, fine-tunes, and implements large language models (LLMs) for various applications, such as chatbots, content generation, and AI-driven tools. They work with machine learning frameworks, optimize model performance, and ensure efficient deployment. This role requires expertise in natural language processing (NLP), deep learning, and programming languages like Python.

What are the key skills and qualifications needed to thrive as an LLM Developer?

To excel as an LLM Developer, you need strong expertise in natural language processing (NLP), deep learning frameworks, and programming languages such as Python, typically supported by a degree in computer science or a related field. Familiarity with machine learning libraries (like TensorFlow or PyTorch), cloud computing platforms, and experience with prompt engineering or fine-tuning large language models is crucial. Excellent problem-solving abilities, collaboration, and effective communication skills help you design solutions and work efficiently within multidisciplinary teams. These qualifications are essential for successfully building, deploying, and optimizing large language models that drive impactful AI applications.

What is the role of a Llm developer?

A Large Language Model (LLM) developer designs, trains, and fine-tunes large-scale AI models for natural language processing tasks. They work with machine learning frameworks, handle large datasets, and optimize models for performance and accuracy, often requiring knowledge of programming languages like Python and tools such as TensorFlow or PyTorch.
What are the most commonly searched types of Llm Developer jobs in California? The most popular types of Llm Developer jobs in California are:
What cities in California are hiring for Llm Developer jobs? Cities in California with the most Llm Developer job openings:
Infographic showing various Llm Developer job openings in California as of August 2026, with employment types broken down into 79% Full Time, 5% Part Time, 2% Temporary, and 14% Contract. Highlights an 80% Physical, 6% Hybrid, and 14% Remote job distribution, with an average salary of $102,977 per year, or $49.5 per hour.

Principal LLM Inference Engineer

Entrada Ventures

Santa Clara, CA • On-site

$120 - $160/hr

Other

Posted 5 days ago


Job description

Role Overview

We are hiring end-to-end inference engineers who are comfortable going from a novel research idea to a deployed, optimized system. You will work at every layer of the inference stack — from kernel-level optimization to distributed orchestration to high-level serving APIs.

This role could be a great match for you if you:
  • Have deep intuition for modern generative AI architectures and how to squeeze performance out of them at inference time.
  • Are familiar with the internals of open-source inference frameworks (vLLM, SGLang, TensorRT-LLM, etc.) and can extend or replace them when needed.
  • Enjoy pathfinding new use cases — exploring heterogeneous deployment topologies and building early-stage POCs that prove out new ideas.
  • Are results-oriented with a strong bias toward action; you own problems end-to-end from prototype to optimization to handoff.
  • Are energized by working at the intersection of novel hardware and frontier models, and want your work to directly influence how next-generation AI silicon is used.
  • Value clear communication and thrive in a small, high-ownership team environment.
Responsibilities
  • Identify and prototype emerging LLM inference use cases suited to heterogeneous hardware deployments.
  • Build compelling proof-of-concept systems that demonstrate D-Matrix capabilities to customers, partners, and internal stakeholders.
  • Develop and tune custom kernels and operator-level optimizations to maximize throughput and minimize latency.
  • Drive quantization, sparsity, and batching strategies tailored to D-Matrix computational model.
  • Build and maintain inference runtimes, serving frameworks, and evaluation tooling.
  • Contribute to distributed inference systems: tensor/pipeline parallelism, disaggregated prefill/decode, KV-cache management.
  • Work closely with hardware architects to provide firmware and compiler teams with actionable inference workload insights.
  • Partner with product and business development to translate POCs into customer-facing demonstrations.
  • Contribute to technical publications, whitepapers, and open-source projects that advance D-Matrix visibility.
Required Qualifications
  • Bachelor’s degree in Computer Science, Electrical Engineering, or a related field, and 10+ years of relevant engineering experience; or equivalent demonstrated experience.
  • Master’s or PhD in Computer Science, Electrical Engineering, or a related field preferred, with 6+ years of relevant industry experience.
  • Strong proficiency in Python and C/C++.
  • Hands‑on experience optimizing LLM inference — attention kernels, KV cache, batching strategies, quantization (INT8/FP8/INT4).
  • Experience with at least one major inference framework (vLLM, SGLang, TensorRT-LLM, ONNX Runtime, or similar) at a contributor level.
  • Familiarity with GPU kernel programming (CUDA/Triton) and performance profiling tools.
Preferred Qualifications
  • Experience with heterogeneous compute deployments — scheduling inference workloads across dissimilar hardware (accelerators, CPUs, GPUs).
  • Familiarity with custom silicon or ASIC-based inference (beyond GPU-only environments).
  • Experience with distributed inference: tensor parallelism, pipeline parallelism, disaggregated serving.
  • Contributions to open-source inference or ML systems projects.
  • Experience with production inference serving at scale (latency SLOs, continuous batching, multi-model serving).
  • Familiarity with speculative decoding, mixture-of-experts routing, or long-context serving techniques.
  • Working familiarity with the material in the JAX Scaling Book or equivalent systems-level understanding of modern LLM training and inference.
Why D-Matrix Frontier Group
  • Work on genuinely novel hardware — D-Matrix in-memory compute architecture opens up inference optimization problems that don’t exist anywhere else.
  • End-to-end ownership from idea to deployed system, with a short feedback loop between your work and real hardware.
  • Small, senior team with high autonomy and direct influence on product direction.
  • Competitive compensation, equity, and benefits in Santa Clara, CA.
Equal Opportunity Employment Policy

d-Matrix is proud to be an equal opportunity workplace and affirmative action employer. We’re committed to fostering an inclusive environment where everyone feels welcomed and empowered to do their best work. We hire the best talent for our teams, regardless of race, religion, color, age, disability, sex, gender identity, sexual orientation, ancestry, genetic information, marital status, national origin, political affiliation, or veteran status. Our focus is on hiring teammates with humble expertise, kindness, dedication and a willingness to embrace challenges and learn together every day.

#J-18808-Ljbffr