1

Evening Machine Learning R Jobs in New York, NY (NOW HIRING)

Machine Learning Engineer, Evals

New York, NY · On-site

$153K/yr

Build a judge calibration protocol: sample human-labeled decisions, measure agreement (κ, per-class P/R), identify drift zones, and document it so anyone can re-run it * Extend an existing benchmark ...

Hands‑on programming skills in SQL, Python, R, or similar languages * Intermediate exposure to machine learning techniques and causal models * Advanced presentation and communication skills

New

Proficiency in at least one programming language (Python or R) and experience using machine learning libraries and tools (e.g., scikit-learn or equivalent R packages) Knowledge of predictive modeling ...

They are seeking a Senior Data Scientist with extensive experience in machine learning algorithms ... on R, Python, IBM SPSS • Excellent communication skills and ability to express the statistical ...

... R • Knowledge of building and applying machine learning or predictive modeling • Strong problem solving skills • Exceptional ability to communicate and present findings clearly to both ...

Proficiency in programming languages such as Python or R.. * Strong understanding of machine learning techniques and algorithms. Preferred Qualifications: * Experience with modern ML frameworks such ...

Proficiency in programming languages such as Python or R.. * Strong understanding of machine learning techniques and algorithms. Preferred Qualifications: * Experience with modern ML frameworks such ...

Proficiency in programming languages such as Python or R.. Strong understanding of machine learning techniques and algorithms. Preferred Qualifications: Experience with modern ML frameworks such as ...

New

AI/ML Engineer

New York, NY · On-site

$175K/yr

... machine learning techniques (e.g., logistic regression, decision trees, neural networks, random forests, etc.) * 2+ years of strong hands-on experience in Python or R * Experience working with both ...

Showing results 41-60

Evening Machine Learning R information

See New York, NY salary details

$27.9K

$46.6K

$96.3K

How much do evening machine learning r jobs pay per year?

As of Aug 18, 2026, the average yearly pay for evening machine learning r in New York, NY is $46,588.00, according to ZipRecruiter salary data. Most workers in this role earn between $35,600.00 and $50,300.00 per year, depending on experience, location, and employer.

What are the most commonly searched types of Machine Learning R jobs in New York, NY?

The most popular types of Machine Learning R jobs in New York, NY are:

Infographic showing various Evening Machine Learning R job openings in New York, NY as of August 2026, with employment types broken down into 1% As Needed, 72% Full Time, 24% Part Time, 1% Temporary, and 2% Contract. Highlights an 87% Physical, 3% Hybrid, and 10% Remote job distribution, with an average salary of $46,588 per year, or $22.4 per hour.

Machine Learning Engineer, Evals

NOUS RESEARCH

New York, NY • On-site

$153K/yr

Full-time

Re-posted 2 days ago


Job description

You'll work across the lab on agent capability evals, benchmark design, LLM-as-judge systems, failure analysis, and the infrastructure that ties it together. This is a high-growth, high-ownership role on a small team, and you'll ship evaluation infrastructure that researchers depend on from day one.
Responsibilities:
  • Run the full eval pipeline end to end and reproduce known results during onboarding, pairing with a senior engineer on your first task
  • Build a judge calibration protocol: sample human-labeled decisions, measure agreement (κ, per-class P/R), identify drift zones, and document it so anyone can re-run it
  • Extend an existing benchmark (GAIA, τ-Bench, SWE-bench slice, etc.) with new tasks targeting known capability gaps, including the prompt, environment, rubric, automated grader, and QA
  • Run failure analysis on model outputs: categorize failure modes, quantify prevalence, and write up findings with recommendations for training data, judge prompts, or benchmark changes
  • Own a recurring eval workflow (weekly regression suite, judge drift dashboard, red-team evaluation for a new capability) and ship tooling researchers actually use

Qualifications:
  • 3+ years in software engineering, ML engineering, data science, or a research-adjacent role, with concrete evaluation experience from coursework, an internship, a side project, open source work, or a job
  • Experience with at least one LLM evaluation framework (Harbor, Nemo Evaluator, etc.), with real opinions on what it does well and where it falls short
  • Hands-on experience with LLMs: prompting, few-shot design, and ideally fine-tuning or RAG; regular use of coding agents
  • Solid Python. You write clean, tested, version-controlled code that a colleague could run without you babysitting it
  • Comfort with Git, CI/CD basics, Docker, and the Linux command line (SSH, tmux, debugging a remote job)
  • Understanding of basic eval statistics: why accuracy misleads on imbalanced judges, what Cohen's κ measures, how to think about confidence intervals on a metric
  • At least 3 of the following: you can explain why LLM-as-judge needs calibration; you've done failure analysis and can tell model bugs apart from prompt, grader, or retrieval issues; you know at least two agent benchmarks (GAIA, AgentBench, τ-Bench, MINT, SWE-bench, WebShop, ALFWorld) and a limitation of each; you've designed or extended an eval dataset with happy paths, edge cases, and adversarial examples; you've thought about non-determinism in eval, how you sample, how many runs, how you report variance
  • You communicate clearly to both researchers and engineers, in the right language for each
  • You're comfortable with ambiguity, can turn a half-formed request into a plan, and know when to ask for help

Preferred:
  • RLVR / RLHF pipeline experience
  • Training data curation experience
  • Distributed eval orchestration experience
  • Benchmark design from scratch
  • Red teaming and adversarial eval experience
  • Familiarity with psychometrics or measurement theory