1

Python Llm Jobs in Round Rock, TX (NOW HIRING)

Deep understanding of LLM data quality challenges and common failure modes. * Experience designing automated tests for AI/ML models. * Familiarity with Python and testing frameworks such as PyTest ...

MTS DevOps Engineer[On-Site}

Manor, TX · On-site

$56.75 - $78/hr

You'll build automation (Ansible/Terraform/Python), scale AWS and Linux systems, and run CI/CD that keeps production fast and stable. You'll also help keep LLM/AI workloads reliable, observable ...

AI Engineer, Data Science

Austin, TX · On-site

$113K - $136K/yr

Commercial experience with modern LLM ecosystems (e.g., LangChain, LlamaIndex, RAG pipelines, multi-agent frameworks, GPT-based systems) * Python experience focused on machine learning and NLP

The Staff Compiler Engineer will lead the development of compilers for their novel LLM accelerator ... Python and C • Experience with hardware compilers • Familiarity with Large Language Model ...

QA

Austin, TX · On-site

$41 - $55.75/hr

... AI/LLM solutions. * Evaluate AI model performance, accuracy, reliability, and safety. * Develop AI testing and evaluation pipelines using Python. * Collaborate with AI/ML teams to improve AI ...

The role requires a strong foundation in Java and Python, as well as experience in developing and ... LLM libraries (e.g., Anthropic SDK) • Understanding of LLM orchestration patterns, prompt ...

QA Engineer - Gen AI

Austin, TX · On-site

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

Deep understanding of LLM data quality challenges and common failure modes. * Experience designing automated tests for AI/ML models. * Familiarity with Python and testing frameworks such as PyTest ...

Support Engineer

Austin, TX

$184K - $277K/yr

  • Medical

  • Dental

  • Retirement

... query issues, and LLM output regressions. Support the full-stack web engineering team by ... Deliver targeted bug fixes and enhancements across the stack - Python microservices, Node.js ...

Deep understanding of LLM data quality challenges and common failure modes. * Experience designing automated tests for AI/ML models. * Familiarity with Python and testing frameworks such as PyTest ...

Support Engineer

Austin, TX

$184K - $277K/yr

  • Medical

  • Dental

  • Retirement

... query issues, and LLM output regressions. Support the full-stack web engineering team by ... Deliver targeted bug fixes and enhancements across the stack - Python microservices, Node.js ...

Showing results 41-60

Python Llm information

See Round Rock, TX salary details

$12

$54

$80

How much do python llm jobs pay per hour?

As of Aug 13, 2026, the average hourly pay for python llm in Round Rock, TX is $54.67, according to ZipRecruiter salary data. Most workers in this role earn between $45.05 and $62.12 per hour, depending on experience, location, and employer.

What is a Python LLM?

A Python LLM job involves working with Large Language Models (LLMs) using Python to develop, fine-tune, and deploy AI models. Responsibilities may include data preprocessing, prompt engineering, model optimization, and integration with applications. Professionals in this role often work with frameworks like TensorFlow, PyTorch, or Hugging Face Transformers. They may also contribute to improving model efficiency, reducing bias, and ensuring ethical AI usage.

What are the key skills and qualifications needed to thrive in the Python LLM position, and why are they important?

To excel as a Python LLM (Large Language Model) Engineer, you need strong skills in Python programming, machine learning, and natural language processing, typically supported by a degree in computer science or a related field. Proficiency with libraries such as TensorFlow, PyTorch, Hugging Face Transformers, and experience with model deployment platforms are often essential, alongside certifications in AI or data science. Effective communication, problem-solving abilities, and collaboration are important soft skills for working in interdisciplinary teams and delivering results in dynamic environments. These skills ensure the development, fine-tuning, and deployment of advanced language models that meet both technical and business objectives.

What are some common challenges faced by Python LLM engineers in their daily work?

Python LLM Engineers often encounter challenges related to optimizing model performance, managing large datasets, and adapting models to specific business needs. Working with large-scale language models requires balancing computational resource limitations with the need for high accuracy and efficiency. Collaboration with data scientists, product managers, and DevOps engineers is routine to ensure seamless model integration and deployment. Staying updated on the latest advancements in NLP and continuously improving models based on user feedback are also important aspects of the role.

What are popular job titles related to Python Llm jobs in Round Rock, TX?

For Python Llm jobs in Round Rock, TX, the most frequently searched job titles are:

What job categories do people searching Python Llm jobs in Round Rock, TX look for?

The top searched job categories for Python Llm jobs in Round Rock, TX are:

What cities near Round Rock, TX are hiring for Python Llm jobs?

Cities near Round Rock, TX with the most Python Llm job openings:

QA Engineer - Gen AI

Sustainment

Austin, TX • On-site

Full-time

Posted 8 days ago


Job description

This is a contract opportunity

Job Overview: We are seeking a QA Engineer to help ensure the reliability, accuracy, and robustness of our AI Agents. This role will focus on data quality, model evaluation, and regression testing frameworks to identify and mitigate common LLM failure modes. You will be responsible for designing automated and scalable quality assurance systems while working in an AWS-based infrastructure. If you have a strong background in LLM testing, data validation, and automated QA frameworks, this role is an excellent opportunity to contribute to cutting-edge AI systems.

Responsibilities: 

  • Design and run regression test suites for LLM evaluation.
  • Identify and track LLM failure modes, including hallucinations, biases, factual inconsistencies, and logical errors.
  • Design data-quality checks to assess training and test datasets.
  • Automate LLM performance monitoring using advanced metrics and validation strategies.
  • Apply best practices for prompt-engineering testing, fine-tuning validation, and output-consistency analysis.
  • Collaborate with ML engineers, data scientists, and product teams to align on quality benchmarks.
  • Work within an AWS ecosystem, leveraging services such as EKS, S3, SageMaker, or Databricks for model testing and evaluation.
  • Build tools and dashboards to track LLM quality over time.
  • Curate and version the ground-truth datasets that serve as the accuracy baseline for document parsing, and translate business and domain requirements into written, testable field definitions (partnering with the labeling team on annotation guidelines).
  • Evaluate structured extraction from real business documents (multi-page PDFs, scans, spreadsheets) by scoring model output field-by-field against ground truth, with tolerance-aware comparison for numbers, dates, free text, and repeated structures.
  • Maintain the ground-truth corpus as a versioned, evolving test asset: keep existing annotations valid as extraction schemas change, preserve dataset provenance, and grow the corpus from real production failures so every customer-reported miss becomes a permanent regression case.
  • Calibrate and validate automated scoring itself; confirm that semantic/LLM-judge scoring agrees with human judgment.

Qualifications:

  • 3+ years in software testing and quality assurance
  • 2+ years with a focus on ML evaluation, NLP, LLMs, VLMs, etc.
  • Deep understanding of LLM data quality challenges and common failure modes.
  • Experience designing automated tests for AI/ML models.
  • Familiarity with Python and testing frameworks such as PyTest, Hypothesis, or similar.
  • Knowledge of evaluation metrics for LLMs (DeepEval, MLflow, LangSmith, or similar).
  • Hands-on experience with automated data validation techniques.
  • Strong debugging and analytical skills.
  • Experience creating or working with labeled evaluation datasets ("golden" sets) for model evaluation.
  • Working knowledge of evaluation metrics for structured information extraction: field-level precision, recall, and F1; exact vs. fuzzy matching; numeric tolerance; and alignment of repeated or nested records.
  • Experience translating ambiguous business requirements into precise, documented field definitions in collaboration with non-technical subject-matter experts.

Preferred Qualifications

  • SQL proficiency, including seeding test data across Postgres environments (local/dev/staging/prod).
  • Comfort with observability and incident-response tooling (e.g., Datadog monitors, alerting/triage) for monitoring and debugging.
  • Familiarity with RAG and RAGAS.
  • Familiarity with containerized dev environments (Kubernetes/Tilt).
  • Understanding of human-in-the-loop (HITL) evaluation strategies.
  • Familiarity with LLM APIs (OpenAI, Anthropic, Bedrock, or similar).
  • Background in statistical analysis or model interpretability.
  • Experience with MLOps practices and CI/CD pipelines for ML models.
  • Experience evaluating document AI / OCR pipelines and their specific failure modes: layout and table extraction, multi-page documents, scanned or low-quality source material.
  • Experience running controlled models and prompt comparison studies.
  • Familiarity with the .NET+Linux ecosystem.
  • Able to read DB schema changes & migrations (EF Core/.NET, DDL).