2

Llm Remote Jobs in Texas (NOW HIRING)

AI Engineer Location: 100% Remote Duration: 6+ month contract-to-hire Requirement: * Implemented ... LLM optimization & improved response handling. Responsibilities: * Design, develop, and deploy ...

... for LLM-based systems. - Ensure semantic accuracy and safety for clinical use cases. Technical ... The starting pay range for this remote role is $105,840.00-$147,000.00. This range reflects the ...

... for LLM-based systems. - Ensure semantic accuracy and safety for clinical use cases. Technical ... The starting pay range for this remote role is $105,840.00-$147,000.00. This range reflects the ...

... for LLM-based systems. - Ensure semantic accuracy and safety for clinical use cases. Technical ... The starting pay range for this remote role is $105,840.00-$147,000.00. This range reflects the ...

$139K - $168K/yr

Experience of transformer models and LLM applications * Strong knowledge of Python or C++, or the ... LI-SS2 LI-REMOTE

$139K - $168K/yr

Experience of transformer models and LLM applications * Strong knowledge of Python or C++, or the ... LI-SS2 LI-REMOTE

$139K - $168K/yr

Experience of transformer models and LLM applications * Strong knowledge of Python or C++, or the ... LI-SS2 LI-REMOTE

$139K - $168K/yr

Experience of transformer models and LLM applications * Strong knowledge of Python or C++, or the ... LI-SS2 LI-REMOTE

$139K - $168K/yr

Experience of transformer models and LLM applications * Strong knowledge of Python or C++, or the ... LI-SS2 LI-REMOTE

$139K - $168K/yr

Experience of transformer models and LLM applications * Strong knowledge of Python or C++, or the ... LI-SS2 LI-REMOTE

$139K - $168K/yr

Experience of transformer models and LLM applications * Strong knowledge of Python or C++, or the ... LI-SS2 LI-REMOTE

$220K - $260K/yr

A key part of our vision is an AI assistant powered by emerging LLM techniques that helps brokers ... Work in a remote-first, flexible, and accountable engineering culture while learning from teammates ...

$200K - $250K/yr

A key part of our vision is an AI assistant powered by emerging LLM techniques that helps brokers ... Work in a remote-first, flexible, and accountable engineering culture while learning from teammates ...

Cybersecurity Engineer

Austin, TX ยท Remote

$123K/yr

A remote position does not require job duties be performed within proximity of a Visa office ... Deep Knowledge of LLM Platforms (Mandatory)Demonstrated handson and architectural knowledge of ...

next page

Showing results 1-20

Llm Remote information

What is an llm remote job?

An LLM Remote job typically refers to a position that involves working with large language models (LLMs) such as OpenAI's GPT, but done remotely rather than in a traditional office setting. These roles can include positions like machine learning engineer, data scientist, prompt engineer, or AI researcher, all focused on developing, fine-tuning, or applying LLMs. Working remotely allows professionals to contribute to AI projects from anywhere, often collaborating with distributed teams and leveraging cloud-based tools. This flexibility is ideal for those who want to work in the AI field without relocating to a tech hub.

What is the difference between Llm Remote vs Legal Assistant?

AspectLlm RemoteLegal Assistant
Required CredentialsLaw degree (JD or equivalent), bar admission (preferred)High school diploma or associate degree, paralegal certification often preferred
Work EnvironmentRemote, flexible hours, legal firms or corporate legal departmentsOffice-based or hybrid, law firms, corporate legal departments
Industry UsageLegal research, document review, legal analysisLegal support, document preparation, client communication
Search & Comparison IntentUnderstanding remote legal roles, legal research jobsLegal support roles, paralegal or legal assistant positions

While both roles support legal operations, Llm Remote typically involves legal research and analysis requiring a law degree, often performed remotely. Legal Assistants focus on administrative and support tasks, usually in-office or hybrid, with less emphasis on legal research. The choice depends on your credentials and preferred work environment.

What are some common challenges faced by remote large language model (LLM) engineers, and how can they overcome them?

Remote LLM engineers often face challenges such as collaborating effectively across time zones, maintaining clear communication with distributed teams, and staying updated on rapidly evolving AI research. To overcome these obstacles, it's important to leverage collaboration tools (like Slack, GitHub, and video conferencing), establish regular check-ins, and participate in virtual knowledge-sharing sessions. Additionally, proactively seeking feedback and engaging with global AI communities can help remote LLM engineers stay aligned with team goals and industry trends.

What are the key skills and qualifications needed to thrive as an llm remote engineer, and why are they important?

To excel as an LLM Remote Engineer, a solid background in machine learning, natural language processing, and proficiency with programming languages like Python is essential, often supported by a degree in computer science or a related field. Experience with frameworks such as PyTorch or TensorFlow, familiarity with large language models (LLMs), and relevant cloud platforms (like AWS or Azure) are typically required, along with certifications in AI or ML being advantageous. Strong problem-solving, communication, and self-motivation are crucial soft skills for collaborating effectively across remote teams and driving innovation. These competencies ensure successful model development, deployment, and maintenance in a distributed work environment.
What are the most commonly searched types of Llm jobs in Texas? The most popular types of Llm jobs in Texas are:
What cities in Texas are hiring for Llm Remote jobs? Cities in Texas with the most Llm Remote job openings:
Infographic showing various Llm Remote job openings in Texas as of August 2026, with employment types broken down into 100% Full Time. Highlights an 100% Remote job distribution.

Applied Data Scientist, LLM Evaluation

Driver AI Inc.

Austin, TX โ€ข Remote

Full-time

Medical, Dental, Vision, Life, Retirement

Re-posted 19 days ago


Job description

Applied Data Scientist, LLM Evaluation Introduction

At Driver, we're building systems that turn source code into human language. The tech stack includes a core compiler-like engine, a heavily asynchronous/distributed backend server, and a frontend web application that provides a rich user experience.

About Driver

We're an early-stage startup backed by Y Combinator and Google Ventures that combines first principles technical approaches and applied LLM expertise to tackle context engineering at scale. Driver builds the context layer for employees and AI agents alike to use in developing software.

Working at Driver

Driver is an early-stage but fast-growing startup. As such, we take advantage of that which startups can excel: delivery speed, flexibility, and enjoying working with a small close-knit team.

Organizational and engineering values at Driver include first-principles thinking, correct by construction, writing things down, experimentation and iteration, pragmatism, commitment to effective communication and transparency, autonomy, and ambition.

Job Overview

Title: Applied Data Scientist, LLM Evaluation

Location: Remote or Austin, Tx

Our value is directly tied to the quality of our content at scale. The platform generates technical documentation across a complex, multi-stage pipeline - producing multiple content types at different levels of abstraction, from individual code elements up to high-level summaries. Today, changes to models, context strategies, or pipeline architecture are evaluated largely through manual review and intuition. There is no systematic way to answer: "Did this change make our output better, worse, or the same - and for which languages, repo sizes, and content types?"

This is a hard problem. LLM outputs are non-deterministic - identical inputs produce different outputs across runs, and small variations at early pipeline stages compound into meaningfully different end-user content downstream. Evaluating quality requires methodology that accounts for this: statistical reasoning over multiple runs, understanding of cascade effects through the pipeline, and rubrics that balance human judgment with automated signals.

This role builds the evaluation function from scratch. You'll define what "good" means for our generated content, build the infrastructure to measure it, and create the experimental framework that lets the team ship changes with confidence.

What You'll Do

You'll own the LLM evaluation strategy at Driver - from first principles to production infrastructure. This is a foundational role: you're not joining an existing eval team, you're building it. As the function matures, you'll seed and grow a team around it.

Define quality metrics and build evaluation datasets. Establish what "good" looks like for each content type across the pipeline. Build and curate gold-standard evaluation datasets across languages and repo archetypes (monorepos, microservices, libraries, applications). Design rubrics that capture accuracy, completeness, usefulness, and readability.

Build benchmarking and experimentation infrastructure. Create automated evaluation pipelines that score output against reference datasets. Instrument the content generation pipeline to support A/B comparisons - run the same codebase through two strategies and compare results. Build tooling for LLM-as-judge evaluation and regression detection. Integrate evaluation into CI so pipeline changes come with quality evidence.

Develop automated quality signals at scale. Build quality checks that flag degraded output without requiring human review of every document. Monitor content quality trends over time. Design sampling strategies for human review that maximize signal with minimal annotation effort.

Quantify tradeoffs and inform decisions. Run experiments on model selection, context strategies, and pipeline architecture changes. Quantify cost/quality/latency tradeoffs. Partner with the engineering team to turn evaluation insights into shipped improvements.

Qualifications

Education: Bachelor's, Master's, or PhD in Statistics, Machine Learning, Data Science, Computational Linguistics, or a related quantitative field.

Experience: Minimum 3 - 5 years in applied science, ML engineering, or data science roles with a focus on evaluation, NLP, or generative AI. 7+ years experience preferred.

Required Technical Skills

  • Strong statistical foundations: experimental design, hypothesis testing, confidence intervals, effect sizes, power analysis.
  • Experience designing and running evaluations for LLM or NLP systems - you've thought carefully about what "better" means when outputs are open-ended text.
  • Proficient in Python and the scientific/data stack (pandas, NumPy, scipy, sklearn).
  • Comfortable working in Jupyter notebooks for exploration and prototyping, and turning that work into automated pipelines.
  • Experience with LLM-as-judge approaches, inter-annotator agreement, and rubric design for subjective quality assessment.
  • Familiarity with the practical challenges of non-deterministic systems: variance decomposition, multi-run methodology, distinguishing signal from noise at scale.
  • Strong data storytelling - you can turn experiment results into clear recommendations that drive engineering and product decisions.

Preferred and Nice-to-Have Technical Skills

  • Experience with LLM APIs and prompt engineering across multiple providers.
  • Familiarity with evaluation frameworks (e.g., RAGAS, DeepEval, custom harnesses).
  • Experience building data pipelines or ETL workflows (Airflow, Dagster, or similar).
  • Comfort with SQL and working directly against production data stores.
  • Experience with visualization tools (Matplotlib, Plotly, Streamlit) for building internal dashboards and reports.
  • Background in code understanding, developer tools, or technical documentation.
  • Experience building or managing annotation pipelines and human evaluation workflows.
Benefits
  • Competitive Compensation Packages - Cash & Equity
  • Flexible Work Culture
  • Unlimited Time Off + 12 Paid Company Holidays
  • Insurance - Health, Dental, & Vision
  • Life Insurance & FSA Accounts
  • 401(k) Retirement Accounts - Traditional, Roth, or Both
  • Quarterly Team Offsites

Driver is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.