1

Huggingface Jobs in California (NOW HIRING)

Software Engineer, LLM Compilation

Cupertino, CA · On-site

$128K - $172K/yr

Responsibilities : • Write an optimized kernel to compute a new attention variant on our hardware • Implement HuggingFace's `CohereForCausalLM` class using Etched's transformer building blocks ...

Experience with ML/LLM libraries such as vLLM, LangChain, PyTorch, and HuggingFace. * Practical experience developing, deploying, and scaling AI Agents in a production environment. * Ability to ...

... HuggingFace), or custom AI services. • Familiarity with vector databases such as pgvector, Pinecone, or Weaviate, and search APIs. • Solid knowledge of RESTful API design and distributed systems ...

next page

Showing results 1-20

Huggingface information

See California salary details

$8

$25

$60

How much do huggingface jobs pay per hour?

As of Aug 18, 2026, the average hourly pay for huggingface in California is $25.79, according to ZipRecruiter salary data. Most workers in this role earn between $14.83 and $30.13 per hour, depending on experience, location, and employer.

What is a Huggingface job?

A Hugging Face job typically refers to a role at Hugging Face, a company specializing in machine learning and natural language processing (NLP). Employees at Hugging Face work on developing and maintaining open-source AI tools, including the popular Transformers library. Roles range from research and engineering to product and community development, often focusing on advancing state-of-the-art AI models.

What does a typical day look like for an engineer working at Hugging Face?

As an engineer at Hugging Face, your day typically involves collaborating with team members to design, develop, and improve state-of-the-art machine learning models and tools, with a strong focus on open-source NLP projects. You’ll participate in code reviews, experiment with new technologies, engage with the community through forums or GitHub, and help support user questions or issues. Expect a fast-paced, collaborative environment where cross-functional teamwork with product managers, researchers, and other engineers is common. The work is project-driven, with plenty of opportunities to contribute ideas, learn from experts, and advance your technical skills.

What are the key skills and qualifications needed to thrive in the Huggingface position, and why are they important?

To thrive in a role at Hugging Face, you typically need strong skills in machine learning, natural language processing (NLP), and software development, supported by a relevant degree in computer science or a related field. Familiarity with frameworks like PyTorch or TensorFlow, plus experience using version control systems such as Git, are often required; open-source contributions and cloud platform knowledge are a plus. Excellent communication, collaborative teamwork, and problem-solving abilities help candidates stand out in this dynamic, innovation-driven environment. These strengths are crucial because they enable individuals to develop high-impact AI tools, work effectively in interdisciplinary teams, and contribute to open-source communities.

What are the most commonly searched types of Huggingface jobs in California?

The most popular types of Huggingface jobs in California are:

What job categories do people searching Huggingface jobs in California look for?

The top searched job categories for Huggingface jobs in California are:

What cities in California are hiring for Huggingface jobs?

Cities in California with the most Huggingface job openings:

Infographic showing various Huggingface job openings in California as of August 2026, with employment types broken down into 99% Full Time, and 1% Contract. Highlights an 83% Physical, 3% Hybrid, and 14% Remote job distribution, with an average salary of $53,653 per year, or $25.8 per hour.

ML Ops Engineer -- Agentic AI Lab (Founding Team)

Fabrion

San Francisco, CA • On-site

Full-time

Re-posted 9 hours ago


Job description

Job Summary:
Fabrion is a pioneering company in intelligent infrastructure through open-source LLMs and AI solutions. They are seeking an ML Ops Engineer to automate model training, deployment, and governance pipelines, ensuring the integration of ML research with production systems.
Responsibilities:
• Build and maintain secure, scalable, and automated pipelines for:
• LLM fine-tuning, SFT, LoRA, RLHF, DPO training
• RAG embedding pipelines with dynamic updates
• Model conversion, quantization, and inference rollout
• Manage hybrid compute infrastructure (cloud, on-prem, GPU clusters) for training and
• inference workloads using Kubernetes, Ray, and Terraform
• Containerize models and agents using Docker, with reproducible builds and CI/CD via
• GitHub Actions or ArgoCD
• Implement and enforce model governance: versioning, metadata, lineage, reproducibility,
• and evaluation capture
• Create and manage evaluation and benchmarking frameworks (e.g. OpenLLM-Evals,
• RAGAS, LangSmith)
• Integrate with security and access control layers (OPA, ABAC, Keycloak) to enforce
• model policies per tenant
• Instrument observability for model latency, token usage, performance metrics, error
• tracing, and drift detection
• Support deployment of agentic apps with LangGraph, LangChain, and custom inference
• backends (e.g. vLLM, TGI, Triton)
Qualifications:
Required:
• 4+ years in MLOps, ML platform engineering, or infra-focused ML roles
• Deep familiarity with model lifecycle management tools: MLflow, Weights & Biases, DVC, HuggingFace Hub
• Experience with large model deployments (open-source LLMs preferred): LLaMA, Mistral, Falcon, Mixtral
• Comfortable with tuning libraries (HuggingFace Trainer, DeepSpeed, FSDP, QLoRA)
• Familiarity with inference serving: vLLM, TGI, Ray Serve, Triton Inference Server
• Proficient with Terraform, Helm, K8s, and container orchestration
• Experience with CI/CD for ML (e.g. GitHub Actions + model checkpoints)
• Managed hybrid workloads across GPU cloud (Lambda, Modal, HuggingFace Inference, Sagemaker)
• Familiar with cost optimization (spot instance scaling, batch prioritization, model sharding)
• Familiarity with LangChain, LangGraph, LlamaIndex or similar RAG/agent orchestration tools
• Built embedding pipelines for multi-source documents (PDF, JSON, CSV, HTML)
• Integrated with vector databases (Weaviate, Qdrant, FAISS, Chroma)
• Implemented model-level RBAC, usage tracking, audit trails
• Integrated with API rate limits, tenant billing, and SLA observability
• Experience with policy-as-code systems (OPA, Rego) and access layers
• 5+ years as a full stack or backend engineer
• Experience owning and delivering production systems end-to-end
• Prior experience with modern frontend frameworks (React, Next.js)
• Familiarity with building APIs, databases, cloud infrastructure, or deployment workflows at scale
• Comfortable working in early-stage startups or autonomous roles, prior experience as a founder, founding engineer, or a 0-1 pre-seed startup is a big plus
Preferred:
• LLM Ops: HuggingFace, DeepSpeed, MLflow, Weights & Biases, DVC
• Infra: Kubernetes (GKE/EKS), Ray, Terraform, Helm, GitHub Actions, ArgoCD
• Serving: vLLM, TGI, Triton, Ray Serve
• Pipelines: Prefect, Airflow, Dagster
• Monitoring: Prometheus, Grafana, OpenTelemetry, LangSmith
• Security: OPA (Rego), Keycloak, Vault
• Languages: Python (primary), Bash, optionally Rust or Go for tooling
• Bonus: experience with SOC2, HIPAA, or GovCloud-grade model operations
• Builder's mindset with startup autonomy: you automate what slows you down
• Obsessive about reproducibility, observability, and traceability
• Comfortable with a hybrid team of AI researchers, DevOps, and backend engineers
• Interested in aligning ML systems to product delivery, not just papers
• Comfortable with ambiguity, eager to prototype and iterate quickly
• Strong sense of ownership — prefers to build systems rather than wait for tickets
• Enjoys thinking about architecture, performance, and tradeoffs at every level
• Clear communicator and pragmatic team player
• Values equity and impact over prestige or hierarchy
• Prior startup or founding team experience
Company:
Fabrion is an AI-native platform purpose-built for the new industrial era Founded in 2025, the company is headquartered in San Francisco, USA, with a team of 2-10 employees. The company is currently Early Stage.