1

Generative Ai Testing Jobs in Boston, MA (NOW HIRING)

Lead the design, development, testing, and deployment of machine learning and artificial ... Manage AI engineering workstreams by assigning work, reviewing deliverables, and driving quality ...

Integrate AI/ML protein design methods with structural biology and high-throughput experimental ... Testing, Therapeutic Proteins, Workflow Optimization Preferred Skills: Current Employees apply HERE ...

Generative AI Strategy Drive the strategic roadmap for generative AI capabilities within the Help ... The Testing Center is fully onboarded in production, with evaluation coverage scaled well beyond ...

New

next page

Showing results 1-20

Generative Ai Testing information

See Boston, MA salary details

$34

$58

$83

How much do generative ai testing jobs pay per hour?

As of Aug 10, 2026, the average hourly pay for generative ai testing in Boston, MA is $58.37, according to ZipRecruiter salary data. Most workers in this role earn between $48.03 and $66.88 per hour, depending on experience, location, and employer.

What is the difference between Generative Ai Testing vs Data Scientist?

AspectGenerative Ai TestingData Scientist
Required CredentialsKnowledge of AI models, testing tools, programming skillsStatistics, programming, data analysis certifications
Work EnvironmentAI development teams, testing labs, tech companiesResearch labs, tech firms, finance, healthcare
Employer & Industry UsageAI product testing, quality assurance in techData analysis, predictive modeling across industries

Generative Ai Testing focuses on evaluating and validating AI-generated content and models, ensuring quality and accuracy. Data Scientists analyze data, build models, and derive insights. While both roles require programming and AI knowledge, Generative Ai Testing emphasizes testing processes, whereas Data Scientists focus on data analysis and model development.

How do I become a Generative AI Testing?

To become a Generative AI Tester, develop skills in machine learning, natural language processing, and programming languages like Python. Gain experience with AI frameworks such as TensorFlow or PyTorch and understand data quality and model evaluation techniques. Certifications in AI or data science can enhance your qualifications and improve job prospects.

Is Generative AI Testing a good career?

Generative AI Testing is a growing field within AI development that involves evaluating the quality and safety of AI-generated content. It requires skills in machine learning, programming, and understanding AI models, making it a promising career path with increasing demand as AI technologies expand. Professionals in this area can find opportunities in tech companies, research labs, and startups focused on AI innovation.

What are the key skills and qualifications needed to thrive as a generative AI testing specialist, and why are they important?

To thrive as a Generative AI Testing Specialist, you need a robust understanding of machine learning principles, model evaluation techniques, and a background in computer science or a related field. Familiarity with tools such as Python, TensorFlow, PyTorch, and model evaluation frameworks, as well as experience with automated testing platforms, is typically required. Analytical thinking, attention to detail, and strong communication skills help you identify model weaknesses and collaborate effectively with development teams. These skills are crucial to ensure the reliability, safety, and ethical deployment of generative AI solutions.

What are some common challenges faced when testing generative AI models, and how can I prepare to address them in this role?

Testing generative AI models often involves unique challenges such as evaluating the quality and relevance of generated content, detecting bias or inappropriate outputs, and ensuring model consistency across various prompts. You may work closely with data scientists and engineers to create robust evaluation frameworks and develop automated as well as manual testing strategies. Familiarity with prompt engineering, statistical evaluation techniques, and domain-specific knowledge will help you address these challenges effectively. Proactively staying updated on industry best practices and collaborating with cross-functional teams are key to success in this dynamic field.

What is generative AI testing?

Generative AI Testing refers to the process of evaluating and validating AI systems, particularly those that generate content such as text, images, or code. This type of testing focuses on assessing the accuracy, reliability, fairness, and safety of generative models to ensure they function as intended and avoid producing harmful or biased outputs. Testers use various methods, including automated and manual techniques, to check for issues like hallucinations, inappropriate content, or security vulnerabilities. The goal is to build trust in generative AI systems and ensure they meet quality and ethical standards before deployment.
What are popular job titles related to Generative Ai Testing jobs in Boston, MA? For Generative Ai Testing jobs in Boston, MA, the most frequently searched job titles are:
What job categories do people searching Generative Ai Testing jobs in Boston, MA look for? The top searched job categories for Generative Ai Testing jobs in Boston, MA are:
What cities near Boston, MA are hiring for Generative Ai Testing jobs? Cities near Boston, MA with the most Generative Ai Testing job openings:
Infographic showing various Generative Ai Testing job openings in Boston, MA as of August 2026, with employment types broken down into 2% Internship, 92% Full Time, 2% Part Time, and 4% Contract. Highlights an 80% In-person, 8% Hybrid, and 12% Remote job distribution, with an average salary of $121,406 per year, or $58.4 per hour.

Sr Data Scientist- Generative AI

Citizens

Westwood, MA

Full-time

Posted 6 days ago


Job description

Description

Join a team where innovation meets impact. As a Senior Data Scientist, Generative AI & Agentic Systems, you will help drive the bank's AI transformation by designing, developing, and deploying Large Language Model (LLM) solutions, Retrieval-Augmented Generation (RAG) systems, AI agents, and intelligent automation capabilities. You will work across business, technology, risk, and compliance teams to deliver responsible, scalable, and production-ready GenAI solutions that improve customer experiences, enhance operational efficiency, and create measurable business value.

This role is ideal for an experienced data scientist with strong software engineering and machine learning skills, deep expertise in NLP and Generative AI, and experience developing AI solutions within highly regulated environments.

Key Responsibilities

  • Design, develop, and deploy production-grade Generative AI solutions using LLMs, RAG frameworks, AI agents, and workflow orchestration platforms.
  • Build intelligent document processing capabilities for information extraction, summarization, classification, question answering, and conversational AI applications.
  • Develop agentic workflows capable of autonomous reasoning, task execution, tool utilization, and multi-step decision support.
  • Design and implement retrieval pipelines, vector search architectures, embedding strategies, and knowledge-grounded AI systems.
  • Evaluate and improve LLM performance through prompt engineering, model benchmarking, hallucination reduction, and faithfulness testing.
  • Build scalable AI solutions using modern frameworks and infrastructure including vLLM, LangChain, LangGraph, MLflow, Databricks, Snowflake, and cloud-native platforms.
  • Perform exploratory data analysis, feature engineering, and statistical analysis to support machine learning and GenAI model development.
  • Develop model monitoring, evaluation, and observability frameworks to measure quality, reliability, fairness, and operational performance.
  • Collaborate closely with Model Risk Management (MRM), Compliance, Audit, Legal, and Information Security teams to ensure responsible AI deployment.
  • Create technical documentation, model development artifacts, validation packages, and executive-level presentations.
  • Partner with product managers, engineers, data architects, and business stakeholders to identify and prioritize GenAI opportunities.
  • Stay current with advances in Generative AI, agentic systems, multimodal AI, foundation models, and emerging industry best practices.

Qualifications

Required

  • Ph.D. or Master's degree in Computer Science, Data Science, Statistics, Mathematics, Artificial Intelligence, or a related quantitative field.
  • 7+ years of experience in data science, machine learning, predictive analytics, or artificial intelligence.
  • 4+ years of hands-on experience developing NLP and Generative AI solutions.
  • Strong proficiency in Python and modern software development practices.
  • Experience developing and deploying LLM-based applications using commercial or open-source models.
  • Experience with Retrieval-Augmented Generation (RAG), vector databases, embeddings, and semantic search.
  • Experience with prompt engineering, prompt evaluation, and LLM performance optimization.
  • Strong understanding of machine learning algorithms, deep learning, statistical modeling, and model explainability techniques.
  • Experience working with structured and unstructured data at enterprise scale.
  • Experience collaborating with cross-functional stakeholders and communicating technical concepts to non-technical audiences.
  • Strong knowledge of model governance, validation processes, and documentation standards.

Preferred

  • Experience designing and deploying AI agents and multi-agent systems.
  • Experience with agent orchestration frameworks such as LangChain, LangGraph, Semantic Kernel, CrewAI, Autogen, or similar technologies.
  • Experience serving open-source LLMs using vLLM, Hugging Face, or equivalent inference frameworks.
  • Experience with RAG evaluation frameworks such as RAGAS or other LLM evaluation methodologies.
  • Experience with model monitoring, MLOps, and production AI deployment.
  • Experience with cloud AI platforms such as AWS Bedrock, Azure AI, Databricks, Snowflake Cortex.
  • Experience building document intelligence solutions involving PDFs, OCR,  document extraction, knowledge extraction from images, and workflow automation.
  • Experience within banking, financial services, fintech, insurance, or other regulated industries.
  • Experience supporting Model Risk Management (MRM), model validation, audit reviews, or regulatory examinations.
  • Familiarity with MCP (Model Context Protocol), tool calling frameworks, and AI workflow automation platforms.

Technical Skills

Generative AI & LLMs

  • GPT, Claude, Llama and other foundation models
  • Retrieval-Augmented Generation (RAG)
  • AI Agents and Multi-Agent Systems
  • Prompt Engineering and Prompt Optimization
  • Fine-Tuning and Model Adaptation
  • LLM Evaluation and Guardrails
  • Knowledge Retrieval and Vector Search

Programming & Frameworks

  • Python
  • SQL
  • PyTorch
  • TensorFlow
  • Scikit-Learn
  • LangChain
  • LangGraph
  • Hugging Face

Data Platforms & MLOps

  • Experience with cloud-based data, AI, and ML platforms (AWS, SageMaker, Databricks, Snowflake, etc.)
  • Experience with distributed data processing frameworks (Spark / PySpark/Snowpark Snowflake)
  • Experience with ML lifecycle, orchestration, and deployment tools (MLflow, Airflow, CI/CD)
  • Experience with AI-assisted development and model monitoring solutions

NLP & Analytics

  • Text Classification
  • Information Extraction
  • Summarization
  • Topic Modeling
  • Question Answering
  • Sentiment Analysis
  • Explainable AI

Preferred Candidate Profile

The ideal candidate needs to demonstrate success building production-scale GenAI solutions such as RAG platforms, conversational AI systems, document intelligence solutions, AI agents, and automated decision-support systems. They possess strong technical depth, understand governance requirements in regulated industries, and can bridge the gap between cutting-edge AI capabilities and practical business outcomes. This individual is comfortable operating from concept through production deployment while maintaining a strong focus on quality, compliance, explainability, and measurable impact.

Hours & Work Schedule

  • Hours per Week: 40
  • Work Schedule: Monday - Friday
  • Hybrid: 4 days per week on-site, 1 day remote

Some job boards have started using jobseeker-reported data to estimate salary ranges for roles. If you apply and qualify for this role, a recruiter will discuss accurate pay guidance.

Equal Employment Opportunity

Citizens, its parent, subsidiaries, and related companies (Citizens) provide equal employment and advancement opportunities to all colleagues and applicants for employment without regard to age, ancestry, color, citizenship, physical or mental disability, perceived disability or history or record of a disability, ethnicity, gender, gender identity or expression, genetic information, genetic characteristic, marital or domestic partner status, victim of domestic violence, family status/parenthood, medical condition, military or veteran status, national origin, pregnancy/childbirth/lactation, colleague's or a dependent's reproductive health decision making, race, religion, sex, sexual orientation, or any other category protected by federal, state and/or local laws. At Citizens, we are committed to fostering an inclusive culture that enables all colleagues to bring their best selves to work every day and everyone is expected to be treated with respect and professionalism. Employment decisions are based solely on merit, qualifications, performance and capability.

Education:Why Work for UsEmployment Type: 1ST