Establish standardized offline and online evaluation methodologies, golden datasets, and quality benchmarks across AI products * Define the roadmap for AI observability, enabling teams to monitor ...
Establish standardized offline and online evaluation methodologies, golden datasets, and quality benchmarks across AI products * Define the roadmap for AI observability, enabling teams to monitor ...
AIML - Sr Lead Engineer, Evaluation
$205K - $374K/yr
We are looking for an A/B experimentation engineering leader to help design and build the next generation of online evaluation platform. You'll work with the teams and technical leaders at the heart ...
AIML - Sr Lead Engineer, Evaluation
$205K - $374K/yr
We are looking for an A/B experimentation engineering leader to help design and build the next generation of online evaluation platform. You'll work with the teams and technical leaders at the heart ...
Independent Contract Clinical Psychologist - KS - 200
Dodge City, KS · On-site
$75K - $104K/yr
Willing to complete R3C's on-line Evaluation and IPE training.
Independent Contract Clinical Psychologist - KS - 200
Dodge City, KS · On-site
$75K - $104K/yr
Willing to complete R3C's on-line Evaluation and IPE training.
AIML - Sr Lead Engineer, Evaluation
Seattle, WA · On-site
$201.30 - $367.40/hr
Seattle, Washington, United States Machine Learning and AI We are looking for an A/B experimentation engineering leader to help design and build the next generation of online evaluation platform. You ...
AIML - Sr Lead Engineer, Evaluation
Seattle, WA · On-site
$201.30 - $367.40/hr
Seattle, Washington, United States Machine Learning and AI We are looking for an A/B experimentation engineering leader to help design and build the next generation of online evaluation platform. You ...
Willing to complete R3C's on-line Evaluation and IPE training
Willing to complete R3C's on-line Evaluation and IPE training
Willing to complete R3C's on-line Evaluation and IPE training.
Willing to complete R3C's on-line Evaluation and IPE training.
Willing to complete R3C's on-line Evaluation and IPE training
Willing to complete R3C's on-line Evaluation and IPE training
Willing to complete R3C's on-line Evaluation and IPE training
Willing to complete R3C's on-line Evaluation and IPE training
Willing to complete R3C's on-line Evaluation and IPE training.
Willing to complete R3C's on-line Evaluation and IPE training.
Willing to complete R3C's on-line Evaluation and IPE training
Willing to complete R3C's on-line Evaluation and IPE training
Willing to complete R3C's on-line Evaluation and IPE training
Willing to complete R3C's on-line Evaluation and IPE training
Willing to complete R3C's on-line Evaluation and IPE training
Willing to complete R3C's on-line Evaluation and IPE training
Willing to complete R3C's on-line Evaluation and IPE training
Willing to complete R3C's on-line Evaluation and IPE training
Willing to complete R3C's on-line Evaluation and IPE training
Willing to complete R3C's on-line Evaluation and IPE training
Willing to complete R3C's on-line Evaluation and IPE training
Willing to complete R3C's on-line Evaluation and IPE training
Psychiatrist - ID - 200
Cascade, ID · On-site
... on-line Evaluation and IPE training.
Psychiatrist - ID - 200
Cascade, ID · On-site
... on-line Evaluation and IPE training.
Senior Machine Learning Engineer, User Signal & Ads
$138K - $182K/yr
Own the model lifecycle : data preparation, feature engineering, training, offline/online evaluation, deployment, monitoring, and iteration - with rigorous A/B testing and clear business metrics (CTR ...
Senior Machine Learning Engineer, User Signal & Ads
$138K - $182K/yr
Own the model lifecycle : data preparation, feature engineering, training, offline/online evaluation, deployment, monitoring, and iteration - with rigorous A/B testing and clear business metrics (CTR ...
Founding Machine Learning Engineer
Boston, MA · On-site
$180 - $240/hr
Stand up offline and online evaluation infrastructure -- measure the gap between them, don't assume it. * Publish ranking and matching APIs for product surfaces, with latency and quality SLOs.
Founding Machine Learning Engineer
Boston, MA · On-site
$180 - $240/hr
Stand up offline and online evaluation infrastructure -- measure the gap between them, don't assume it. * Publish ranking and matching APIs for product surfaces, with latency and quality SLOs.
Stand up offline and online evaluation infrastructure - measure the gap between them, don't assume it. * Publish ranking and matching APIs for product surfaces, with latency and quality SLOs.
Stand up offline and online evaluation infrastructure - measure the gap between them, don't assume it. * Publish ranking and matching APIs for product surfaces, with latency and quality SLOs.
... online evaluation infrastructure -- measure the gap between them, don't assume it. • Publish ranking and matching APIs for product surfaces, with latency and quality SLOs. • Instrument model ...
... online evaluation infrastructure -- measure the gap between them, don't assume it. • Publish ranking and matching APIs for product surfaces, with latency and quality SLOs. • Instrument model ...
Online Evaluation information
See salary details
$17.5K - $23.7K
12% of jobs
$28.2K is the 25th percentile. Wages below this are outliers.
$23.7K - $30K
18% of jobs
$30K - $36.2K
15% of jobs
The median wage is $37.1K / yr.
$36.2K - $42.4K
33% of jobs
$42.4K - $48.6K
12% of jobs
$48.6K - $54.9K
0% of jobs
$54.9K - $61.1K
0% of jobs
$61.1K - $67.3K
0% of jobs
$67.3K - $73.5K
0% of jobs
$73.5K - $79.8K
0% of jobs
$79.8K - $86K
9% of jobs
$17.5K
$40.6K
$86K
How much do online evaluation jobs pay per year?
What are the key skills and qualifications needed to thrive as an online evaluator, and why are they important?
What is online evaluation?
What is the difference between Online Evaluation vs Quality Assurance Specialist?
| Aspect | Online Evaluation | Quality Assurance Specialist |
|---|---|---|
| Credentials | Typically requires a degree in education, psychology, or related fields | Requires a degree in quality management, engineering, or related fields |
| Work Environment | Primarily remote, online platforms, educational or testing organizations | Office settings, manufacturing, or software development environments |
| Industry Usage | Used in education, testing, and e-learning sectors | Common in manufacturing, software, and product development industries |
| Job Focus | Assessing and scoring online tests or educational content | Ensuring product quality and process improvements |
Online Evaluation and Quality Assurance Specialist roles share a focus on assessment and quality, but differ in industry, credentials, and work environment. Online Evaluation is centered on educational and testing contexts, often remote, while Quality Assurance Specialists work across industries to improve product quality in more physical or technical settings.
What are some common challenges faced by online evaluators, and how can they effectively overcome them?

Full-time
PTO
Posted 4 days ago
Job description
As a Product Manager on Ridgeline's AI Platform team, you'll define and execute the strategy for the systems that ensure our AI capabilities are accurate, reliable, and continuously improving. You'll lead initiatives spanning LLM evaluation frameworks, AI model benchmarking, evaluation harnesses, prompt and model experimentation, AI quality platforms, and observability capabilities that help teams understand, measure, and improve production AI performance. You'll work closely with engineering, AI researchers, design, and product teams to deliver platform capabilities that enable every AI-powered experience across Ridgeline. You'll leverage cutting-edge AI technologies and tools in a fast-moving, creative, progressive work environment while helping establish best practices for developing and operating enterprise AI products.
At Ridgeline, how we work matters as much as what we build. Ridgeliners act like owners, choose growth over comfort, and communicate with transparency. We assume positive intent, bias toward action, and bring solutions-not just problems. We celebrate wins, learn from setbacks, and thrive in a resilient, collaborative, high-performing culture. If this excites you, we'd love to meet you!
You must be work authorized in the United States without the need for employer sponsorship.
The impact you will have
- Define and execute the product strategy and roadmap for Ridgeline's AI Quality Platform
- Build scalable LLM evaluation frameworks and automated evaluation harnesses that enable continuous model validation
- Establish standardized offline and online evaluation methodologies, golden datasets, and quality benchmarks across AI products
- Define the roadmap for AI observability, enabling teams to monitor quality, latency, cost, reliability, and user outcomes in production
- Partner with engineering and AI researchers to build continuous regression testing and experimentation capabilities for generative AI systems
- Drive product decisions using AI quality metrics, customer feedback, and quantitative evaluation data
- Collaborate with engineering teams to deliver platform capabilities that enable trusted, measurable, and continuously improving AI experiences across Ridgeline
- Translate advances in foundation models, evaluation techniques, and AI tooling into platform capabilities that create meaningful customer value
- Prioritize product investments using customer insights, product analytics, and measurable business outcomes
- Communicate product vision, roadmap, and priorities clearly across technical and executive stakeholders
- Foster a collaborative environment rooted in ownership, continuous learning, experimentation, and transparency
What we look for
- 5+ years of product management experience building highly technical platform, infrastructure, or developer-focused products
- Demonstrated experience shipping production AI products powered by large language models
- Experience defining products that incorporate LLM evaluation frameworks, AI quality measurement, and continuous model evaluation
- Deep understanding of offline and online evaluation methodologies, model benchmarking, regression testing, and AI quality metrics
- Experience with LLM observability platforms such as Datadog LLM Observability, Langfuse, LangSmith, Arize AI, Braintrust, Galileo, or similar AI quality and evaluation solutions
- Familiarity with prompt engineering, experimentation frameworks, evaluation datasets, and LLM-as-a-judge techniques for generative AI applications
- Experience partnering effectively with software engineers, ML engineers, AI researchers, and cross-functional stakeholders
- Strong analytical skills with the ability to translate AI performance signals into product decisions
- Excellent written and verbal communication skills
- Willingness to learn about cutting-edge technologies while cultivating expertise in investment management workflows and enterprise AI
- An aptitude for problem solving
- Ability to communicate effectively
- Serious interest in having fun at work
Bonus
- Experience building products around retrieval-augmented generation (RAG), AI agents, MCP, or model orchestration frameworks
- Experience with synthetic data generation and automated evaluation pipelines
- Experience building internal developer platforms or ML infrastructure products
- Experience in enterprise SaaS, developer platforms, or financial technology
About Ridgeline
Ridgeline is headquartered in Lake Tahoe, with offices in New York, Reno, the Bay Area, Dublin Ireland. Ridgeline is recognized by Fast Company as a "Best Workplace for Innovators," by Frost & Sullivan as a "Technology Innovation Leader," and by The Software Report as a "Top 100 Software Company."
Ridgeline is proud to be a community-minded, discrimination-free equal opportunity workplace.
Ridgeline processes the information you submit in connection with your application in accordance with the Ridgeline Applicant Privacy Statement. Please review the Ridgeline Applicant Privacy Statement in full to understand our privacy practices and contact us with any questions.
Compensation and Benefits
The typical starting salary range for new hires in this role is listed below. In select locations (including, the San Francisco Bay Area, CA, and the New York City Metro Area), an alternate range may apply as specified below.
The typical starting salary range for this role is: $136,000-$170,000.
The typical starting salary range for this role in the select locations listed above is: $149,500-$187,000.
Final compensation amounts are determined by multiple factors, including candidate experience and expertise, and may vary from the amount listed above.
As an employee at Ridgeline, you'll have many opportunities for advancement in your career and can make a true impact on the product.
In addition to the base salary, Ridgeline employees can participate in our Company Stock Plan subject to the applicable Stock Option Agreement. We also offer rich benefits that reflect the kind of organization we want to be: one in which our employees feel valued and are inspired to bring their best selves to work. These include unlimited vacation, educational and wellness reimbursements, and $0 cost employee insurance plans. Please check out our Careers page for a more comprehensive overview of our perks and benefits.