2

Remote Large Language Model Llm Jobs (NOW HIRING)

AI Architect

OR · Remote

Artificial and Large Language Model Architect As an AI & LLM Architect , you will play a pivotal role in designing and implementing the technology architecture for advanced AI (including Large ...

... Large Language Model (LLM) solution to check coherence and consistency. Resource will also develop and modify existing models related to customer long term engagement and retention. This role will ...

Hands-on experience with large language model (LLM) inference and/or model training using open ... Remote-friendly within the United States * Preference for candidates located near major East or ...

... remote within a mutually acceptable location. #LI-Hybrid Success Looks Like: * AI systems move ... Design and implement AI-powered applications including large language model (LLM) systems and ...

next page

Showing results 1-20

Remote Large Language Model Llm information

See salary details

$14

$24

$38

How much do remote large language model llm jobs pay per hour?

As of Aug 4, 2026, the average hourly pay for remote large language model llm in the United States is $24.34, according to ZipRecruiter salary data. Most workers in this role earn between $18.99 and $29.09 per hour, depending on experience, location, and employer.

What is a remote large language model LLM?

A Remote Large Language Model (LLM) job involves working with advanced AI models, like GPT or similar, from a remote location. Professionals in these roles may develop, train, fine-tune, or implement large language models for various applications such as natural language processing, chatbots, or content generation. Remote LLM jobs can include positions like machine learning engineer, research scientist, or AI product manager. The work typically requires strong programming skills, experience with AI frameworks, and the ability to collaborate virtually with global teams.

How does a remote large language model LLM engineer typically collaborate with cross-functional teams while working remotely?

Remote LLM Engineers often work closely with data scientists, product managers, and software engineers through virtual meetings, collaborative coding platforms, and shared documentation tools. Regular communication is key, with daily stand-ups or weekly syncs to align on project goals, update progress, and address challenges. They may also participate in code reviews, contribute to design discussions, and support model deployment efforts, all within a distributed team environment. This remote structure encourages self-motivation and proactive communication to ensure project success.

What are the key skills and qualifications needed to thrive as a remote large language model LLM engineer?

To thrive as a Remote Large Language Model (LLM) Engineer, you need a strong background in computer science, machine learning, and natural language processing, typically supported by a relevant degree and experience with large-scale models. Proficiency with programming languages like Python, deep learning frameworks such as PyTorch or TensorFlow, and familiarity with cloud platforms and distributed systems are essential. Excellent problem-solving, communication, and collaboration skills are critical for remote teamwork and translating complex requirements into scalable solutions. These skills ensure the effective development, deployment, and maintenance of advanced language models in fast-evolving, distributed environments.

What is the difference between Remote Large Language Model Llm vs Data Scientist?

AspectRemote Large Language Model LlmData Scientist
Required CredentialsAdvanced degrees in AI, NLP, or related fields; experience with machine learning frameworksDegree in Data Science, Statistics, Computer Science, or related fields; strong analytical skills
Work EnvironmentPrimarily remote, focused on developing and fine-tuning language modelsRemote or on-site, analyzing data, building models, and generating insights
Employer & Industry UsageTech companies, AI research labs, startups working on NLP productsTech firms, finance, healthcare, marketing, and research organizations

While both roles involve data and machine learning, a Remote Large Language Model Llm specializes in developing and refining language models, whereas a Data Scientist focuses on analyzing data, building predictive models, and deriving insights across various domains.

More about Remote Large Language Model Llm jobs
What cities are hiring for Remote Large Language Model Llm jobs? Cities with the most Remote Large Language Model Llm job openings:
What are the most commonly searched types of Large Language Model Llm jobs? The most popular types of Large Language Model Llm jobs are:
What states have the most Remote Large Language Model Llm jobs? States with the most job openings for Remote Large Language Model Llm jobs include:
Infographic showing various Remote Large Language Model Llm job openings in the United States as of July 2026, with employment types broken down into 1% As Needed, 71% Full Time, 23% Part Time, 1% Temporary, and 4% Contract. Highlights an 93% Physical, 1% Hybrid, and 6% Remote job distribution, with an average salary of $50,625 per year, or $24.3 per hour.

Natural Language Measurement Specialist

College Board

Remote

Full-time

Posted 8 days ago


Job description

Natural Language Specialist, Educational Measurement & AI
College Board - Learning & Assessment
Location:
  • This is a remote role. Candidates who live near CB offices have the option of being fully remote or hybrid (Tuesday and Wednesday in office). All CB employees are required to occasionally travel to meet in person for business purposes.

Role Type:
  • This is a full-time position

About the Team
The Automated Scoring team provides critical insights and tools to support the design, delivery, and continuous improvement of digital assessments. We operate at the intersection of educational measurement, data science, and emerging AI technologies. Our work spans the measurement of language-based constructs, the development of large language model (LLM) systems for feedback generation and annotation, and research to ensure the validity, fairness, and reliability of our systems.
We are a collaborative, mission-driven team that values both psychometric rigor and technical expertise. We combine modern machine learning approaches, including large language models, with strong measurement principles to create scalable, trustworthy solutions that expand opportunity for students.
About the Opportunity
As a Natural Language Specialist, you will help define and advance how language-based performance is measured in high-stakes educational settings. This role sits at the intersection of natural language processing, large language models, and educational measurement, and is ideal for someone who pairs strong measurement training with a working knowledge of modern AI.
You will translate measurement constructs into LLM-based feedback and annotation systems, design and conduct the studies that establish their validity, fairness, and reliability, and ensure that what our models produce holds up to rigorous psychometric standards. You will serve as the measurement authority for cross-functional partners-psychometricians, engineers, and data scientists-shaping how language-based constructs are defined, evaluated, and applied across the team's portfolio of systems. Your primary lens will be measurement: defining what feedback and annotations mean, evidencing that they mean it, and improving them over time.
In this role, you will:
Natural Language Measurement & Psychometrics (40%)
  • Define and operationalize language-based constructs for automated annotation and feedback generation

  • Apply psychometric principles-reliability, validity, dimensionality, and measurement invariance-to LLM-based feedback and annotation systems

  • Design and lead validity studies, including human-machine agreement, rater comparison, and fairness analyses across subgroups

  • Develop and apply methods for detecting and mitigating bias in language-based scores

  • Establish/Recommend annotation guidelines, feedback quality criteria, and standards for acceptable model performance

  • Translate measurement requirements into specifications that guide model development and evaluation

LLM & AI Development (30%)
  • Contribute to prompt design, fine-tuning, and evaluation of LLM-based feedback and annotation systems

  • Develop and refine machine learning models for measuring language-based constructs

  • Build evaluation frameworks that connect model behavior to measurement outcomes

  • Collaborate with senior team members to translate measurement findings into production systems

  • Implement high-quality, maintainable code for model development and evaluation

Research & Validation (10%)
  • Lead and contribute to research studies that evaluate model performance and support assessment validity

  • Apply statistical and psychometric methods to analyze results and inform model improvements

  • Document methodologies and findings in a clear and rigorous manner

  • Stay current with advances in educational measurement, NLP, and learning science

Data Engineering & Pipelines (10%)
  • Prepare and curate datasets for measurement studies and model evaluation

  • Support reproducible data processing workflows for training, evaluation, and monitoring

  • Partner with engineers to integrate feedback and annotation models into scalable systems

Team Operations & Collaboration (10%)
  • Collaborate closely with psychometricians, data scientists, and engineers

  • Contribute to documentation, methodological standards, and team best practices

  • Participate in peer reviews and knowledge sharing

  • Actively raise the measurement literacy of the broader team-mentoring junior ICs, providing technical feedback on colleagues' work, and building shared standards

About You
You bring strong measurement training and a genuine interest in how modern AI can be used to measure language-based performance. You are excited about applying psychometric rigor to large language models in high-stakes educational settings.
You have:
  • A Master's or PhD (or near completion) in a quantitative field such as Psychometrics, Educational Measurement, Quantitative Psychology, Statistics, Data Science, or a related discipline (measurement-focused training strongly valued)

  • A solid foundation in measurement theory, including reliability, validity, and fairness; familiarity with IRT, generalizability theory, or related frameworks is highly desirable

  • Experience analyzing language or text data, with exposure to NLP or large language models

  • Programming skills in Python and familiarity with data science libraries (e.g., pandas, NumPy, PyTorch)

  • Demonstrated ability to conduct rigorous research (e.g., thesis, publications, or applied research projects)

  • Strong skills in statistical analysis and experimental design

  • Familiarity with working with structured and unstructured data

  • Strong attention to detail and a commitment to producing high-quality, reproducible work

  • The ability to travel 5-10 times a year to College Board offices or on behalf of College Board business.

You are:
  • Curious and eager to learn new tools, methods, and domains

  • Thoughtful about the implications of AI systems, including fairness and validity

  • Able to communicate technical and measurement concepts clearly to diverse audiences

  • Comfortable working in a collaborative, cross-functional environment

  • Motivated by mission-driven work in education

All roles at College Board require:
  • A passion for expanding educational and career opportunities and mission-driven work

  • Curiosity and enthusiasm for emerging technologies, with a willingness to experiment with and adopt new AI-driven solutions and comfort with learning and applying new digital tools independently and proactively.

  • Clear and concise communication skills, written and verbal

  • A learner's mindset and a commitment to growth: welcoming diverse perspectives, giving and receiving timely, respectful feedback, and continuously improving through iterative learning and user input.

  • A drive for impact and excellence: solving complex problems, making data-informed decisions, prioritizing what matters most, and continuously improving through learning, user input, and external benchmarking.

  • A collaborative and empathetic approach: working across differences, fostering trust, and contributing to a culture of shared success

  • Authorization to work in the United States

About Our Process
  • Application review will begin immediately and will continue until the position is filled. This role is expected to accept applications for a minimum of 5 business days.

  • While the hiring process may vary, it generally includes: resume and application submission, recruiter phone/video screen, hiring manager interview, performance exercise such as live coding, a panel interview, a conversation with leadership and reference checks.

What We Offer
At College Board, we offer more than a paycheck- we provide a meaningful career, a supportive team, and a comprehensive package designed to help you thrive. We're a self-sustaining nonprofit that believes in fair and competitive compensation grounded in your qualifications, experience, impact, and the market.
A Thoughtful Approach to Compensation
  • The hiring range for this role is $88,000-$145,000.

  • Your exact salary will depend on your location, experience, and how your background compares to others in similar roles at the College Board.

  • We aim to make our best offer upfront, rooted in fairness, transparency, and market data.

  • We adjust salaries by location to ensure fairness, no matter where you live.

You'll have open, transparent conversations about compensation, benefits, and what it's like to work at College Board throughout your hiring process. Check out our careers page for more.