1

Retrieval Augmented Generation Jobs in Austin, TX

Senior Machine Learning Engineer

Austin, TX · On-site

$103K - $142K/yr

... Retrieval Augmented Generation (RAG) pipelines, including vector database integration for contextual retrieval. • Work with multi-modal AI systems across computer vision, audio, and natural ...

This role defines end-to-end AI solution patterns involving large language models, APIs, retrieval-augmented generation, vector search, intelligent agents, orchestration workflows, Snowflake, cloud ...

... Retrieval-Augmented Generation (RAG) systems to improve AI-generated outputs. • Implement and integrate automated red-teaming solutions like PyRIT, Garak, and Giskard. • Collaborate with cross ...

Senior Machine Learning Engineer

Austin, TX · On-site

$103K - $142K/yr

... Retrieval Augmented Generation (RAG) pipelines, including vector database integration for contextual retrieval. • Work with multi-modal AI systems across computer vision, audio, and natural ...

Senior Machine Learning Engineer

Austin, TX · On-site

$103K - $142K/yr

... Retrieval Augmented Generation (RAG) pipelines, including vector database integration for contextual retrieval. • Work with multi-modal AI systems across computer vision, audio, and natural ...

GenAI Physical Synthesis Engineer

Austin, TX · On-site

$134K - $138K/yr

... RAG (Retrieval-Augmented Generation), and fine-tuning techniques Experience with agentic AI frameworks beyond MCP (AutoGen, CrewAI, LangChain agents, etc.) Background in CAD flow or frontend ...

... retrieval-augmented generation (RAG). • Strong understanding of NLP, deep learning, and model fine-tuning techniques. • Experience working with MLOps, cloud-based AI deployment (AWS/GCP/Azure ...

PyTorch, TensorFlow, CUDA, Jupyter Notebooks, Large Language Models (LLMs) and Open-Source Models (Llama, Anthropic, Mistral), LangChain, Retrieval-Augmented Generation (RAG), Hugging Face, Exo Labs ...

AI Agent Engineer Designs and develops AI-driven agentic solutions, including autonomous workflows and Retrieval-Augmented Generation (RAG) systems, to enhance productivity, automate processes, and ...

PyTorch, TensorFlow, CUDA, Jupyter Notebooks, Large Language Models (LLMs) and Open-Source Models (Llama, Anthropic, Mistral), LangChain, Retrieval-Augmented Generation (RAG), Hugging Face, Exo Labs ...

Experience with large language models, including fine-tuning, retrieval augmented generation, and prompt engineering * Ability to coordinate and drive initiatives with little to no oversight

Showing results 21-40

Retrieval Augmented Generation information

What is a retrieval augmented generation?

A Retrieval Augmented Generation (RAG) job typically involves developing and optimizing AI systems that enhance text generation by incorporating external knowledge retrieved from relevant sources. Professionals in this field work on integrating retrieval mechanisms with large language models to improve the relevance, accuracy, and factual grounding of generated content. Common responsibilities include designing retrieval systems, fine-tuning language models, optimizing performance, and ensuring the seamless integration of factual data into AI-generated text. This role is highly interdisciplinary, involving expertise in natural language processing (NLP), machine learning, and information retrieval.

What does a retrieval augmented generation engineer do?

A Retrieval Augmented Generation engineer typically spends their day designing and implementing systems that combine information retrieval with advanced generative models, such as large language models. This includes fine-tuning models, integrating external data sources, developing vector search pipelines, and evaluating output quality. Collaboration with data scientists, machine learning engineers, and product teams is common to ensure the solutions meet user requirements and scale effectively. Additionally, RAG engineers often troubleshoot issues, monitor model performance in production, and stay informed about the latest advancements in AI and information retrieval.

What skills and qualifications are needed for retrieval augmented generation?

To thrive in a Retrieval Augmented Generation (RAG) engineering role, you need a solid background in machine learning, natural language processing (NLP), and experience with scalable information retrieval systems, typically supported by a relevant degree in computer science or a related field. Familiarity with tools such as Python, PyTorch or TensorFlow, vector databases, and search platforms like Elasticsearch is essential, along with practical experience deploying and tuning RAG pipelines. Strong problem-solving skills, a collaborative mindset, and effective communication abilities set outstanding professionals apart in this field. These competencies are crucial for designing, implementing, and optimizing hybrid retrieval-generation AI systems that address complex, real-world information needs.

What are popular job titles related to Retrieval Augmented Generation jobs in Austin, TX?

For Retrieval Augmented Generation jobs in Austin, TX, the most frequently searched job titles are:

What job categories do people searching Retrieval Augmented Generation jobs in Austin, TX look for?

The top searched job categories for Retrieval Augmented Generation jobs in Austin, TX are:

What cities near Austin, TX are hiring for Retrieval Augmented Generation jobs?

Cities near Austin, TX with the most Retrieval Augmented Generation job openings:

Infographic showing various Retrieval Augmented Generation job openings in Austin, TX as of August 2026, with employment types broken down into 70% Full Time, 28% Part Time, and 2% Contract. Highlights an 71% Physical, 2% Hybrid, and 27% Remote job distribution.

Senior Machine Learning Engineer

webAI

Austin, TX • On-site

$103K - $142K/yr

Full-time

Re-posted 16 days ago


Job description

Job Summary:
webAI is seeking a Senior Machine Learning Engineer to support their Public Sector initiatives focused on building and optimizing production-ready AI systems. The role involves transforming prototype models into scalable and reliable production systems that operate across various hardware environments.
Responsibilities:
• Design, develop, and deploy agentic workflows to orchestrate multi-step reasoning, tool use, and decision-making across production systems.
• Productionize AI models from research prototypes into scalable, deployable systems used in real world applications.
• Engineer adaptive ML systems using LoRA, PEFT, and on-device inference strategies, leveraging PyTorch, TensorFlow, and Hugging Face Transformers for model development, fine-tuning, and optimization.
• Implement model optimization techniques such as quantization, pruning, distillation, and hardware specific acceleration.
• Build and maintain Retrieval Augmented Generation (RAG) pipelines, including vector database integration for contextual retrieval.
• Work with multi-modal AI systems across computer vision, audio, and natural language domains.
• Optimize model execution for distributed and resource constrained environments, ensuring reliability under variable connectivity conditions.
Qualifications:
Required:
• Active US Security clearance
• 4+ years of experience in applied AI, ML engineering, or production AI systems.
• Deep proficiency in PyTorch, TensorFlow, or Hugging Face Transformers.
• Proven experience deploying AI models across cloud, edge, and mobile hardware environments.
• Expertise in model compression and optimization (quantization, pruning, distillation).
• Experience building RAG pipelines and integrating vector databases (e.g., Quadrant, ChromaDB, FAISS, Milvus, Pinecone).
• Familiarity with multi-modal models and synthetic data generation methods.
• Strong algorithmic and problem solving skills, especially in distributed or constrained compute environments.
Preferred:
• Experience with edge AI, federated learning, or offline inference systems.
• Understanding of AI governance and compliance frameworks relevant to public sector deployments.
• Experience integrating models into large scale distributed systems or microservice architectures.
• Excellent communication and technical documentation skills for collaboration across multi disciplinary teams.
• Strong understanding of GPU computing, CUDA, and performance profiling.
Company:
The leader in private AI. Founded in 2020, the company is headquartered in Austin, USA, with a team of 201-500 employees. The company is currently Growth Stage.