1

Llm Backend Engineer Jobs (NOW HIRING)

About the Role Join our fast-growing startup as a Staff Backend Engineer and help scale the ... AI Development Tools: AI-powered coding assistants (Claude Code, Cursor, etc.), LLM-based ...

$100K - $185K/yr

Develop and maintain LLM-based systems capable of extracting insights, automating workflows, and ... modern backend engineering practices. * Lead system design discussions, identify technical ...

Job#: 3040458 Back-End Engineer Location: Dearborn, Michigan (Friday's Remote) Role Overview In ... This position operates within a cloud-native environment, utilizing modern AI and LLM tools to ...

... LLM expertise to tackle context engineering at scale. Driver builds the context layer for employees ... Software Engineer - Backend Location : Remote or Austin, TX About the Role Our core innovation, the ...

We are seeking an experienced Python Backend Engineer to design, develop, and maintain scalable ... Knowledge of AI/LLM integration frameworks such as LangChain or LlamaIndex. Strong experience with ...

Software Engineer - Backend

Austin, TX · On-site +1

$175K - $275K/yr

... LLM expertise to tackle context engineering at scale. Driver builds the context layer for employees ... Software Engineer - Backend Location : Remote or Austin, TX About the Role Our core innovation, the ...

We are seeking a Senior Backend Engineer with strong experience in Functional Java , SpringBoot ... Familiarity with LLM concepts (prompt engineering, APIs for model access, basic NLP tasks)

Associate Backend Engineer

New York, NY · On-site

$95K - $115K/yr

The Role We're looking for a Backend Developer with 1-3 years of experience to join our engineering ... Use AI agents and LLM tooling to prototype, iterate, and ship faster * Contribute product ideas and ...

... LLM responses, batch matching jobs, payment processing, and real-time chat - while maintaining ... Design and build scalable backend services in Node.js/TypeScript, including RESTful APIs ...

The Role We're hiring a Backend Product Engineer to build and maintain the systems that make Tolan ... Contribute to LLM engineering work that brings Tolan's AI capabilities to life. * Work cross ...

They are seeking a Senior Backend Engineer to own and evolve the real-time infrastructure behind ... LLM pipelines deployed in production • Clear communication skills; can translate technical ...

next page

Showing results 1-20

Llm Backend Engineer information

See salary details

$60.5K

$147.7K

$199K

How much do llm backend engineer jobs pay per year?

As of Jul 22, 2026, the average yearly pay for llm backend engineer in the United States is $147,662.00, according to ZipRecruiter salary data. Most workers in this role earn between $124,000.00 and $172,000.00 per year, depending on experience, location, and employer.

How much do LLM engineers make?

LLM backend engineers typically earn between $100,000 and $180,000 annually, depending on experience, location, and company size. Senior roles or those with specialized skills in machine learning frameworks and large language models can command higher salaries, often exceeding $200,000.

What are some common challenges faced by LLM Backend Engineers when deploying large language models in production?

LLM Backend Engineers often encounter challenges such as optimizing inference latency, managing high resource consumption, and ensuring scalability for production workloads. Balancing model performance with cost efficiency requires careful selection of hardware, batching strategies, and model quantization techniques. Additionally, they must address security and privacy concerns associated with handling sensitive data processed by the models. Collaboration with data scientists and DevOps teams is essential to streamline model updates and monitor system health.

What are LLM Backend Engineers?

LLM Backend Engineers are software engineers who specialize in designing, building, and optimizing the backend infrastructure that supports large language models (LLMs) like GPT-4. They focus on integrating LLMs into products and services, ensuring scalable APIs, managing data pipelines, and optimizing inference performance. Their work often involves deploying models in cloud environments, monitoring system reliability, and collaborating with AI researchers to bring advancements into production. LLM Backend Engineers play a critical role in making AI-powered applications robust, efficient, and accessible to end users.

Are LLM engineers in demand?

LLM backend engineers are in high demand due to the rapid growth of artificial intelligence and natural language processing applications. Companies seek professionals skilled in machine learning frameworks, programming languages like Python, and large language model deployment to develop and optimize AI systems. This demand is expected to continue as AI technology becomes more integrated into various industries.

What are the key skills and qualifications needed to thrive as an LLM Backend Engineer, and why are they important?

To thrive as an LLM Backend Engineer, you need a solid foundation in software engineering, backend architecture, and experience working with large language models, typically supported by a degree in computer science or a related field. Proficiency with programming languages like Python or Java, cloud platforms (AWS, GCP, Azure), and machine learning frameworks such as TensorFlow or PyTorch is essential, along with familiarity with APIs and containerization tools like Docker or Kubernetes. Strong problem-solving, collaboration, and communication skills distinguish top performers in this role. These skills ensure robust, scalable, and efficient deployment of LLM-powered applications while enabling effective teamwork and innovation.

What engineer makes 500,000 a year?

Senior engineers in specialized fields such as software engineering, data engineering, or machine learning engineering can earn $500,000 or more annually, especially with extensive experience, advanced skills, and working at large tech companies or in high-demand industries. These roles often require expertise in cloud platforms, programming, and system architecture, along with leadership responsibilities or stock options that contribute to total compensation.

What engineers make $300,000 a year?

Senior machine learning engineers, especially those working with large language models (LLMs) and possessing expertise in deep learning, distributed systems, and cloud infrastructure, can earn $300,000 or more annually. High compensation is often associated with experience, specialized skills, and working at top tech companies or in high-demand industries.
More about Llm Backend Engineer jobs
What cities are hiring for Llm Backend Engineer jobs? Cities with the most Llm Backend Engineer job openings:
What states have the most Llm Backend Engineer jobs? States with the most job openings for Llm Backend Engineer jobs include:
Infographic showing various Llm Backend Engineer job openings in the United States as of July 2026, with employment types broken down into 1% As Needed, 92% Full Time, 2% Part Time, and 5% Contract. Highlights an 79% Physical, 4% Hybrid, and 17% Remote job distribution, with an average salary of $147,662 per year, or $71 per hour.

Senior Backend Engineer

Judgment Labs

San Francisco, CA • On-site

Full-time

Posted 6 days ago


Job description

Senior Backend Engineer
San Francisco • On Site • Full Time
Judgment Labs is building the infrastructure for continual learning in long-horizon AI agents.
The next generation of agents will not improve from prompts alone. They will improve from experience: the tasks they attempt, the tools they use, the mistakes they make, the edge cases they encounter, and the outcomes they produce in production. The hard part is turning that raw experience into high-quality data that can actually improve the system.
Judgment builds the infrastructure to do that. We turn long agent trajectories into clean, structured data for evals, labeling, rubric generation, context engineering, and RL workflows. Instead of only showing teams what happened, Judgment helps decide what matters, what should be learned from, and how that learning should flow back into the agent.
Databricks built the data infrastructure for analytics. Judgment is building the learning infrastructure for agents.
We've raised $30M+ from Lightspeed, SV Angel, Valor Equity Partners, and others.
The Role
We're looking for a Senior Backend Engineer to own the systems that ingest, structure, evaluate, and serve agent experience data at production scale.
This role includes the backend and data infrastructure surface area: high-throughput telemetry ingestion, ClickHouse-backed OLAP performance, evaluation pipelines, RabbitMQ/Temporal workflows, multi-tenant scheduling, and product-facing APIs. Some weeks you'll be deep in distributed systems and query performance. Other weeks you'll ship a customer-facing feature end to end across backend, frontend, and the data layer.
This is not a narrow API role. The backend is where raw agent trajectories become structured learning data.
Interesting Technical Challenges
  • High-throughput telemetry ingestion. Parse and persist OTEL traces across protobuf and JSON formats at hundreds of thousands of spans per second, writing to ClickHouse while keeping ingest latency low and backpressure graceful as customer traffic spikes.
  • Petabyte-scale OLAP performance. Design schemas, partitioning, indexes, storage layouts, and query paths so behavioral queries over billions of spans stay fast. Turn real access-pattern analysis into concrete data modeling decisions.
  • Long-horizon trajectory modeling. Agent workflows are messy: multi-step tasks, tool calls, retries, partial failures, context changes, and unclear outcomes. Build the abstractions that turn those trajectories into structured data for evals, labeling, rubric generation, context engineering, and RL workflows.
  • Queue- and workflow-driven evaluation. Evaluations fan out across RabbitMQ and Temporal workflows. Getting this right means reasoning about retries, timeouts, idempotent state transitions, exactly-once-ish semantics, and reconciling runs that fail partway so nothing is silently orphaned.
  • Multi-tenant fairness at scale. A single large customer should not be able to starve everyone else. Build scheduling and execution systems so latency stays predictable across hundreds of teams sharing the same evaluation pipeline.
  • Near-real-time scoring. Behavioral scorers and agent judges call LLM APIs at scale, so batching, rate-limit management, retry/backoff, failure handling, and cost control are core backend systems problems.
  • Learning loops for agents. Build the product and systems layer that helps teams decide what matters, what should be learned from, and how that learning flows back into the agent.
What You'll Do
  • Design and build backend systems for trace ingestion, trajectory processing, evaluation orchestration, scoring, labeling, rubric generation, and customer-facing analytics.
  • Own the API surface used by the Judgment platform UI, SDKs, JudgmentHub libraries, MCP server, Slack agent, and customer integrations.
  • Build and operate the RabbitMQ / Temporal evaluation pipeline, including retry semantics, failure recovery, state reconciliation, and tenant-level scheduling.
  • Optimize the ClickHouse OLAP layer: schema design, partitioning, skip indexes, full-text-search pruning, query rewrites, deduplication, pagination correctness, and storage growth.
  • Turn raw spans, conversations, tool calls, scorer outputs, and agent-judge results into clean data models customers can use for evals, labeling, context engineering, and RL workflows.
  • Ship features end to end, often across Next.js, backend APIs, queues/workflows, and the data layer.
  • Work directly with customers to understand where their agents fail, what data is useful, and how Judgment should structure that experience for learning.
  • Roll out safely with feature flags, design docs, code reviews, tests, observability, and production debugging.
  • Raise the engineering bar through clear interfaces, maintainable systems, thoughtful reviews, and strong ownership.
What We're Looking For
  • Strong backend engineering experience building and operating production systems under real load.
  • Excellent fundamentals in distributed systems, API design, data modeling, reliability, and performance.
  • Experience working with high-volume event, trace, log, metric, or telemetry data.
  • Strong intuition for data systems: query patterns, storage layout, indexing, partitioning, latency, correctness, and cost.
  • Comfort owning systems beyond initial launch: debugging production issues, improving observability, scaling bottlenecks, and cleaning up abstractions as the product evolves.
  • Ability to work across backend, data, product, and infrastructure boundaries rather than treating them as separate silos.
  • Product judgment and willingness to ship across the stack when needed.
  • Clear communication. You can write a design doc, review a diff, explain a tradeoff, and unblock others without turning everything into process.
Nice to Have
  • Experience with ClickHouse, OLAP systems, distributed query engines, or large-scale analytical databases.
  • Experience with RabbitMQ, Temporal, Kafka, Spark, Flink, Ray, Airflow, Dagster, Prefect, or similar queue/workflow/data systems.
  • Experience with OTEL, observability products, tracing, logging, or monitoring infrastructure.
  • Experience building systems that call LLM APIs at scale, including rate-limit management, retries, batching, and cost control.
  • Experience with LLM evaluation, labeling systems, rubric generation, context engineering, RL data pipelines, embedding pipelines, vector search, clustering, or anomaly detection.
  • Experience building developer-facing products, SDK-backed platforms, or customer-facing infrastructure.
Why Judgment?
  • We're building the learning infrastructure for agents. As agents move from demos to production, the bottleneck is no longer just better prompts. It is turning real production experience into high-quality data for evals, labeling, rubric generation, context engineering, and RL workflows.
  • The technical problems are foundational. Long agent trajectories are messy, high-volume, and hard to reason about. We're building the systems that ingest them, structure them, evaluate them, surface what matters, and feed that learning back into the agent.
  • This is a Databricks-scale infrastructure opportunity. Databricks built the data infrastructure for analytics. Judgment is building the learning infrastructure for agents.
  • You'll work on problems customers actually feel. Engineers talk directly to teams building production agents, see where their systems fail, and turn those failures into product and infrastructure.
  • Small team, high ownership. You will not own a narrow slice. You'll shape core systems early, ship quickly, and work across product, data, backend, infra, and customer environments.
  • In person in San Francisco. We work together in person because the problems are hard, the product is moving fast, and the feedback loops matter.