1

Deepspeed Jobs in Texas (NOW HIRING)

Train and fine-tune LLMs using PyTorch, DeepSpeed, and LoRA. * Optimize inference using ONNX, vLLM, TensorRT, and GPU acceleration. * Manage datasets, preprocess data, and implement RAG with vector ...

Senior AI Model Fine-Tuning Engineer

Austin, TX ยท On-site

$103K - $142K/yr

... DeepSpeed). โ€ข Experience in evaluating model performance, including using metrics like BLEU, ROUGE, perplexity, and custom evaluation frameworks. โ€ข Candidates must be willing and able to work on ...

Profile and eliminate system-level bottlenecks across the entire AI pipeline, tuning everything from deep learning frameworks (PyTorch, DeepSpeed, etc.) down to OS-level NUMA pinning and I/O ...

Staff Machine Learning Engineer

Austin, TX ยท On-site +1

$208K - $255K/yr

Familiarity with NVIDIA NeMo, Kaldi, ESPnet, Hugging Face, Whisper, DeepSpeed, or equivalent ecosystems. * Strong Python engineering skills and experience building production ML systems. * Experience ...

Experience with distributed training frameworks (DeepSpeed, FairScale, Horovod) and large-scale model training * Familiarity with edge AI deployment and model optimization techniques (quantization ...

Profile and eliminate system-level bottlenecks across the entire AI pipeline, tuning everything from deep learning frameworks (PyTorch, DeepSpeed, etc.) down to OS-level NUMA pinning and I/O ...

Profile and eliminate system-level bottlenecks across the entire AI pipeline, tuning everything from deep learning frameworks (PyTorch, DeepSpeed, etc.) down to OS-level NUMA pinning and I/O ...

Deepspeed information

What are some common challenges faced by engineers working with DeepSpeed and how can they be addressed?

Engineers working with DeepSpeed often encounter challenges related to optimizing large-scale model training, such as managing memory efficiency and tuning distributed training parameters. Troubleshooting issues like gradient accumulation, parallelism strategies, and ensuring compatibility with different hardware setups can be complex. Collaborating closely with data scientists, DevOps, and research teams is essential for addressing these challenges, as is staying updated with the latest DeepSpeed releases and documentation. Regular participation in code reviews and knowledge-sharing sessions can also help engineers overcome technical hurdles and continuously improve model performance.

What is Deepspeed?

Deepspeed is an open-source deep learning optimization library developed by Microsoft, designed to enable distributed training of large-scale models efficiently. It helps researchers and engineers train models that are too large to fit in the memory of a single GPU by offering features like ZeRO optimization, mixed-precision training, and advanced parallelism techniques. Deepspeed is widely used in the machine learning community for its scalability and performance improvements, making it easier to train state-of-the-art models on vast datasets. The library integrates seamlessly with PyTorch and supports training on multiple GPUs and even across multiple machines.

What is the difference between Deepspeed vs Data Scientist?

AspectDeepspeedData Scientist
Required credentialsKnowledge of machine learning frameworks, programming skills in Python, experience with AI model trainingDegree in Data Science, Statistics, Computer Science, or related fields; strong analytical skills
Work environmentAI research labs, tech companies, cloud computing environmentsBusiness, tech companies, research institutions
Industry usageAI model training, deep learning optimizationData analysis, predictive modeling, business insights

Deepspeed focuses on optimizing large-scale AI model training and deep learning performance, while Data Scientists analyze data to generate insights and build predictive models. Both roles require technical skills but serve different purposes within the AI and data ecosystem.

What are the key skills and qualifications needed to thrive as a DeepSpeed Engineer, and why are they important?

To thrive as a DeepSpeed Engineer, you need a solid background in machine learning, deep learning frameworks (such as PyTorch), and distributed systems, often supported by a degree in computer science or a related field. Proficiency with DeepSpeed, parallel computing libraries, and cloud platforms, along with familiarity with tools like CUDA and NCCL, is typically expected. Strong problem-solving abilities, collaboration, and adaptability are crucial soft skills for optimizing large-scale AI models and working with cross-functional teams. Mastering these skills ensures efficient development and deployment of high-performance, scalable AI solutions in demanding environments.
Infographic showing various Deepspeed job openings in Texas as of July 2026, with employment types broken down into 2% Internship, 97% Full Time, and 1% Contract. Highlights an 84% Physical, 3% Hybrid, and 13% Remote job distribution.

GenAI Engineer-W2

Vkore Solutions

Austin, TX โ€ข On-site

Contractor

Re-posted 24 days ago


Job description

We are looking for a GenAI Ops Engineer to train, fine-tune, and deploy Generative AI models (LLMs, Diffusion Models, Transformers, etc.). You will optimize model performance, manage training pipelines, and integrate AI solutions into production.

Key Responsibilities:

  • Train and fine-tune LLMs using PyTorch, DeepSpeed, and LoRA.
  • Optimize inference using ONNX, vLLM, TensorRT, and GPU acceleration.
  • Manage datasets, preprocess data, and implement RAG with vector databases (FAISS, Chroma, Pinecone).
  • Automate training workflows using ML flow, Weights & Biases, and Ray.
  • Deploy models using Kubernetes, Docker, and cloud AI services AWS or GCP.
  • Monitor model performance, mitigate drift, and optimize resource utilization.

Requirements:

  • Experience with LLM training, fine-tuning, and inference optimization.
  • Proficiency in Python, cloud AI services, and distributed training.
  • Familiarity with retrieval-augmented generation (RAG) and prompt engineering.
  • Strong problem-solving skills and ability to work in fast-paced AI environments.

Preferred:

  • Experience with open-weight models (LLaMA, Mistral, Gemma, Falcon, etc.).
  • Hands-on knowledge of multi-agent architectures and synthetic data generation.