1

Deep Learning Quantization Jobs in Austin, TX (NOW HIRING)

Senior AI Infrastructure Engineer

Austin, TX ยท On-site

$107K - $146K/yr

This role requires deep expertise in distributed systems, cloud-native infrastructure, AI platform ... Proven ability to design, build, and operate production AI or machine learning infrastructure.

Staff AI Infrastructure Engineer

Austin, TX ยท On-site

$106K - $139K/yr

This role requires deep expertise in distributed systems, cloud-native infrastructure, AI platform ... Proven ability to design, build, and operate production AI or machine learning infrastructure.

Senior AI Infrastructure Engineer

Austin, TX ยท On-site

$107K - $146K/yr

This role requires deep expertise in distributed systems, cloud-native infrastructure, AI platform ... Proven ability to design, build, and operate production AI or machine learning infrastructure.

Staff AI Infrastructure Engineer

Austin, TX ยท On-site

$106K - $139K/yr

This role requires deep expertise in distributed systems, cloud-native infrastructure, AI platform ... Proven ability to design, build, and operate production AI or machine learning infrastructure.

Showing results 21-24

Deep Learning Quantization information

See Austin, TX salary details

$10.9K

$83.1K

$138.8K

How much do deep learning quantization jobs pay per year?

As of Aug 7, 2026, the average yearly pay for deep learning quantization in Austin, TX is $83,148.00, according to ZipRecruiter salary data. Most workers in this role earn between $71,400.00 and $137,800.00 per year, depending on experience, location, and employer.

What are the key skills and qualifications needed to thrive as a deep learning quantization engineer, and why are they important?

To excel as a Deep Learning Quantization Engineer, you need a strong background in machine learning, applied mathematics, and computer science, usually supported by an advanced degree in a related field. Familiarity with deep learning frameworks (such as TensorFlow or PyTorch), quantization toolkits, and hardware acceleration platforms is crucial. Analytical thinking, problem-solving, and clear technical communication are standout soft skills in this role. These abilities are essential for efficiently optimizing models for deployment on resource-constrained hardware while maintaining accuracy and performance.

What is the difference between Deep Learning Quantization vs Machine Learning Engineer?

AspectDeep Learning QuantizationMachine Learning Engineer
Required CredentialsAdvanced degrees in AI, Computer Science, or related fields; knowledge of neural networksBachelor's or Master's in CS, Data Science, or related fields; programming skills
Work EnvironmentResearch labs, AI development teams, hardware optimization settingsSoftware development teams, data-driven projects, product-focused environments
Industry UsageAI hardware optimization, model deployment, edge computingModel development, data analysis, software solutions across industries

Deep Learning Quantization focuses on reducing model size and improving inference speed through techniques like weight and activation quantization, often in hardware or embedded systems. Machine Learning Engineers develop, implement, and optimize machine learning models for various applications. While both roles require knowledge of AI and programming, Deep Learning Quantization is more specialized in model optimization techniques, whereas Machine Learning Engineers work broadly on model development and deployment.

What is deep learning quantization?

Deep learning quantization is the process of reducing the precision of the numbers used to represent a neural network's parameters, activations, or both. By converting the typically used 32-bit floating-point values to lower bit-width formats such as 16-bit or 8-bit integers, quantization significantly reduces the memory footprint and computational requirements of deep learning models. This technique helps deploy models efficiently on edge devices and mobile hardware while maintaining acceptable accuracy levels. Quantization is widely used in model optimization for faster inference and lower power consumption.

What are some common challenges faced when implementing deep learning quantization in production environments?

One of the main challenges in implementing deep learning quantization is balancing model accuracy with computational efficiency, as quantization can sometimes lead to a drop in model performance. Additionally, ensuring hardware compatibility and optimizing for different devices (such as CPUs, GPUs, or edge devices) can require extensive testing and tuning. Collaboration with data scientists, software engineers, and hardware specialists is often essential to successfully deploy quantized models at scale. Staying updated with the latest quantization techniques and frameworks is also important for overcoming these challenges.
What cities near Austin, TX are hiring for Deep Learning Quantization jobs? Cities near Austin, TX with the most Deep Learning Quantization job openings:
Infographic showing various Deep Learning Quantization job openings in Austin, TX as of August 2026, with employment types broken down into 1% As Needed, 73% Full Time, 23% Part Time, 1% Temporary, and 2% Contract. Highlights an 87% Physical, 2% Hybrid, and 11% Remote job distribution, with an average salary of $83,148 per year, or $40 per hour.

Senior AI Infrastructure Engineer

Seekr

Austin, TX โ€ข On-site

$107K - $146K/yr

Other

Posted 29 days ago


Job description

Seekr is building the infrastructure that powers the next generation of enterprise AI. As a Senior AI Infrastructure Engineer, you will design, build, and operate the platforms that enable large-scale training, serving, evaluation, and deployment of foundation models and autonomous AI agents.

You will work across distributed systems, Kubernetes, GPU infrastructure, high-performance inference, and enterprise AI platforms to build secure, scalable, and highly reliable systems capable of serving workloads ranging from edge AI deployments to trillion-parameter foundation models.

This role requires deep expertise in distributed systems, cloud-native infrastructure, AI platform engineering, and production software development. You will collaborate with research scientists, software engineers, product teams, and infrastructure engineers to define the architecture and technical direction of Seekr's AI platform.

Duties and Responsibilities

  • Design, develop, deploy, and maintain production AI infrastructure supporting model training, fine-tuning, inference, evaluation, and agentic AI workloads.
  • Design and operate scalable Kubernetes-based infrastructure supporting GPU-accelerated workloads across cloud, on-premises, hybrid, and edge environments.
  • Architect and optimize high-performance inference platforms capable of serving models ranging from resource-constrained edge deployments to trillion-parameter foundation models, with a focus on latency, throughput, scalability, reliability, and cost efficiency.
  • Build and maintain distributed systems that enable reliable scheduling, orchestration, deployment, monitoring, and lifecycle management of AI workloads.
  • Develop enterprise platforms supporting autonomous and multi-agent AI systems, including secure tool execution, orchestration, memory, evaluation, governance, and observability.
  • Design, implement, and automate AI infrastructure using Infrastructure-as-Code, GitOps, CI/CD pipelines, and modern software engineering practices.
  • Evaluate and integrate emerging AI infrastructure technologies, model serving frameworks, hardware accelerators, and cloud-native platforms to improve platform performance, scalability, and reliability.
  • Collaborate with engineering, research, product, and cross-functional teams to deliver secure, scalable, and production-ready AI platforms.
  • Lead technical design discussions, perform architecture reviews, mentor engineers, and establish engineering standards and best practices across the AI Infrastructure organization.
  • Participate in production support activities, including troubleshooting complex distributed systems, performance tuning, incident response, and continuous operational improvement.

Skills and Qualifications

  • 5-8 years of professional software engineering experience building distributed systems, cloud infrastructure, or large-scale platform services
  • Strong production ML infra experience, executes complex work independently, owns significant components
  • 4 year or higher degree or additional relevant experience, in addition to years of work experience
  • Demonstrated success designing and operating production Kubernetes environments supporting cloud-native applications and distributed services.
  • Strong software engineering skills using Python and one or more modern programming languages such as Go, Rust, or C++.
  • Proven ability to design, build, and operate production AI or machine learning infrastructure.
  • Expertise developing and optimizing large-scale AI inference platforms, including GPU utilization, distributed inference, batching, caching, quantization, and accelerator performance.
  • Familiarity with modern AI serving technologies such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Ray Serve, or similar platforms.
  • Knowledge of distributed computing, networking, storage systems, cloud-native architectures, and infrastructure automation using technologies such as Kubernetes, Helm, Argo CD, Docker, Prometheus, Grafana, OpenTelemetry, and Infrastructure-as-Code tools.
  • Experience developing enterprise AI platforms, autonomous agents, or multi-agent systems, including orchestration, tool execution, governance, observability, and evaluation.
  • Familiarity with event-driven architectures, distributed messaging systems, and public cloud platforms including AWS, Azure, Oracle Cloud Infrastructure, or Google Cloud Platform.
  • Demonstrated technical leadership, including driving architectural decisions, mentoring engineers, and leading complex technical initiatives across cross-functional teams.
  • Demonstrated ability to analyze, profile, and optimize AI systems for performance, scalability, reliability, and cost across distributed compute environments.