2

Remote Manual Testing Jobs in Austin, TX (NOW HIRING)

Showing results 41-43

Remote Manual Testing information

See Austin, TX salary details

$10

$41

$60

How much do remote manual testing jobs pay per hour?

As of Aug 17, 2026, the average hourly pay for remote manual testing in Austin, TX is $41.54, according to ZipRecruiter salary data. Most workers in this role earn between $33.37 and $48.37 per hour, depending on experience, location, and employer.

What is remote manual testing?

Remote manual testing is a software quality assurance process where testers manually execute test cases and report bugs from a location outside of a traditional office environment, typically from home or another remote location. Unlike automated testing, manual testing requires human intervention to verify application functionality, usability, and identify issues that automated scripts might miss. Remote manual testers use collaboration tools, test management platforms, and communication channels to coordinate with development teams and stakeholders. This approach allows for flexibility in work location while ensuring software meets quality standards before release.

What are the key skills and qualifications needed to thrive as a remote manual tester?

To thrive as a Remote Manual Tester, you need a solid understanding of software testing methodologies, attention to detail, and experience with test case creation and execution, often supported by a relevant degree or QA certification. Familiarity with bug tracking systems like Jira, test management tools such as TestRail, and basic knowledge of databases or web technologies is commonly required. Strong communication, self-motivation, and problem-solving skills are essential for collaborating with distributed teams and managing tasks independently. These skills ensure accurate defect identification, effective teamwork, and high-quality software releases in a remote work environment.

How does collaboration typically work for remote manual testers within distributed teams?

Remote manual testers often collaborate closely with developers, product managers, and automation testers through virtual meetings, project management tools, and communication platforms like Slack or Microsoft Teams. Regular stand-ups and sprint planning sessions help ensure everyone is aligned on testing priorities and timelines. Effective documentation and clear bug reporting are essential, as remote environments rely heavily on written communication to track issues and progress. Being proactive in asking questions and sharing updates helps remote testers stay connected and contribute to team goals.

What is the difference between Remote Manual Testing vs Remote Automation Testing?

AspectRemote Manual TestingRemote Automation Testing
Skills RequiredTest case execution, attention to detail, basic scripting knowledgeScripting, programming, automation tools proficiency
Work EnvironmentPrimarily individual, manual test executionRequires setup of automation frameworks, coding environment
CertificationsISTQB, ISTQB Agile, or similar testing certificationsISTQB, Certified Automation Engineer, or related certifications
Industry UsageUsed across software development, QA teamsUsed in continuous integration, DevOps, large-scale testing

Remote Manual Testing involves executing test cases manually to identify bugs, requiring attention to detail and basic scripting knowledge. Remote Automation Testing focuses on creating and maintaining automated test scripts using specialized tools. Both roles often require similar certifications and are used across various industries, but automation testing is more technical and suited for large-scale, repetitive testing tasks.

Are remote manual testing testers still in demand?

Remote manual testing testers continue to be in demand as companies seek to ensure software quality across various platforms. Strong attention to detail, understanding of testing tools, and communication skills are valuable in this role, which often involves collaborating with development teams remotely. The demand persists due to ongoing software updates and the need for thorough testing before release.

Is remote manual testing in demand?

Remote manual testing is in steady demand as companies seek flexible testing options to ensure software quality. Skills in test case creation, defect reporting, and familiarity with testing tools are valuable, and many organizations are hiring for remote testing roles due to the increasing reliance on digital products.

What are popular job titles related to Remote Manual Testing jobs in Austin, TX?

For Remote Manual Testing jobs in Austin, TX, the most frequently searched job titles are:

What job categories do people searching Remote Manual Testing jobs in Austin, TX look for?

The top searched job categories for Remote Manual Testing jobs in Austin, TX are:

What cities near Austin, TX are hiring for Remote Manual Testing jobs?

Cities near Austin, TX with the most Remote Manual Testing job openings:

Infographic showing various Remote Manual Testing job openings in Austin, TX as of August 2026, with employment types broken down into 68% Full Time, and 32% Part Time. Highlights an 100% Remote job distribution, with an average salary of $86,407 per year, or $41.5 per hour.

Applied Data Scientist, LLM Evaluation

Driver AI Inc.

Austin, TX • On-site, Remote

$175K - $275K/yr

Full-time

Medical, Dental, Vision, Life, Retirement

Re-posted 24 days ago


Job description

Applied Data Scientist, LLM Evaluation
Introduction
At Driver, we're building systems that turn source code into human language. The tech stack includes a core compiler-like engine, a heavily asynchronous/distributed backend server, and a frontend web application that provides a rich user experience.
About Driver
We're an early-stage startup backed by Y Combinator and Google Ventures that combines first principles technical approaches and applied LLM expertise to tackle context engineering at scale. Driver builds the context layer for employees and AI agents alike to use in developing software.
Working at Driver
Driver is an early-stage but fast-growing startup. As such, we take advantage of that which startups can excel: delivery speed, flexibility, and enjoying working with a small close-knit team.
Organizational and engineering values at Driver include first-principles thinking, correct by construction, writing things down, experimentation and iteration, pragmatism, commitment to effective communication and transparency, autonomy, and ambition.
Job Overview
Title: Applied Data Scientist, LLM Evaluation
Location: Remote or Austin, Tx
Our value is directly tied to the quality of our content at scale. The platform generates technical documentation across a complex, multi-stage pipeline - producing multiple content types at different levels of abstraction, from individual code elements up to high-level summaries. Today, changes to models, context strategies, or pipeline architecture are evaluated largely through manual review and intuition. There is no systematic way to answer: "Did this change make our output better, worse, or the same - and for which languages, repo sizes, and content types?"
This is a hard problem. LLM outputs are non-deterministic - identical inputs produce different outputs across runs, and small variations at early pipeline stages compound into meaningfully different end-user content downstream. Evaluating quality requires methodology that accounts for this: statistical reasoning over multiple runs, understanding of cascade effects through the pipeline, and rubrics that balance human judgment with automated signals.
This role builds the evaluation function from scratch. You'll define what "good" means for our generated content, build the infrastructure to measure it, and create the experimental framework that lets the team ship changes with confidence.
What You'll Do
You'll own the LLM evaluation strategy at Driver - from first principles to production infrastructure. This is a foundational role: you're not joining an existing eval team, you're building it. As the function matures, you'll seed and grow a team around it.
Define quality metrics and build evaluation datasets. Establish what "good" looks like for each content type across the pipeline. Build and curate gold-standard evaluation datasets across languages and repo archetypes (monorepos, microservices, libraries, applications). Design rubrics that capture accuracy, completeness, usefulness, and readability.
Build benchmarking and experimentation infrastructure. Create automated evaluation pipelines that score output against reference datasets. Instrument the content generation pipeline to support A/B comparisons - run the same codebase through two strategies and compare results. Build tooling for LLM-as-judge evaluation and regression detection. Integrate evaluation into CI so pipeline changes come with quality evidence.
Develop automated quality signals at scale. Build quality checks that flag degraded output without requiring human review of every document. Monitor content quality trends over time. Design sampling strategies for human review that maximize signal with minimal annotation effort.
Quantify tradeoffs and inform decisions. Run experiments on model selection, context strategies, and pipeline architecture changes. Quantify cost/quality/latency tradeoffs. Partner with the engineering team to turn evaluation insights into shipped improvements.
Qualifications
Education: Bachelor's, Master's, or PhD in Statistics, Machine Learning, Data Science, Computational Linguistics, or a related quantitative field.
Experience: Minimum 3 - 5 years in applied science, ML engineering, or data science roles with a focus on evaluation, NLP, or generative AI. 7+ years experience preferred.
Required Technical Skills
  • Strong statistical foundations: experimental design, hypothesis testing, confidence intervals, effect sizes, power analysis.
  • Experience designing and running evaluations for LLM or NLP systems - you've thought carefully about what "better" means when outputs are open-ended text.
  • Proficient in Python and the scientific/data stack (pandas, NumPy, scipy, sklearn).
  • Comfortable working in Jupyter notebooks for exploration and prototyping, and turning that work into automated pipelines.
  • Experience with LLM-as-judge approaches, inter-annotator agreement, and rubric design for subjective quality assessment.
  • Familiarity with the practical challenges of non-deterministic systems: variance decomposition, multi-run methodology, distinguishing signal from noise at scale.
  • Strong data storytelling - you can turn experiment results into clear recommendations that drive engineering and product decisions.

Preferred and Nice-to-Have Technical Skills
  • Experience with LLM APIs and prompt engineering across multiple providers.
  • Familiarity with evaluation frameworks (e.g., RAGAS, DeepEval, custom harnesses).
  • Experience building data pipelines or ETL workflows (Airflow, Dagster, or similar).
  • Comfort with SQL and working directly against production data stores.
  • Experience with visualization tools (Matplotlib, Plotly, Streamlit) for building internal dashboards and reports.
  • Background in code understanding, developer tools, or technical documentation.
  • Experience building or managing annotation pipelines and human evaluation workflows.
Benefits
  • Competitive Compensation Packages - Cash & Equity
  • Flexible Work Culture
  • Unlimited Time Off + 12 Paid Company Holidays
  • Insurance - Health, Dental, & Vision
  • Life Insurance & FSA Accounts
  • 401(k) Retirement Accounts - Traditional, Roth, or Both
  • Quarterly Team Offsites

Driver is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.