About the job Remote | LLM Training & Alignment Research Scientist - $95-$115/hour
We are sharing a specialised part-time consulting opportunity for experienced machine learning researchers with hands-on expertise in foundation model pre-training, large-scale data pipelines, language model post-training, and empirical LLM research.
This role focuses on well-scoped, open-ended research problems involving the end-to-end training and improvement of transformer-based language models. Selected researchers will train models from scratch, fine-tune open-weight systems, build pre-training corpora and post-training pipelines, diagnose training failures, and investigate methods for improving performance under limited data and compute budgets.
Key Responsibilities
Foundation Model Pre-Training
- Train transformer-based language models from scratch across full end-to-end workflows
- Design experiments involving model size, token allocation, training duration, and compute budgets
- Investigate performance in data- and compute-constrained regimes
- Diagnose optimisation failures, convergence issues, and training instabilities
- Evaluate interventions using rigorous empirical comparisons
Pre-Training Data Development
- Construct training corpora from raw web crawls and other large-scale unfiltered sources
- Develop pipelines for filtering, deduplication, quality classification, and data selection
- Optimise dataset mixtures, sequencing, and curriculum strategies
- Measure the impact of data interventions on downstream model behaviour
- Identify contamination, duplication, quality, and coverage issues within training datasets
LLM Post-Training & Alignment
- Build supervised fine-tuning pipelines using curated, synthetic, weakly supervised, or rejection-sampled datasets
- Conduct preference optimisation using methods such as DPO, RLHF, or RLAIF
- Develop reward models and systems for predicting human preferences
- Improve refusal behaviour, truthfulness, robustness, and unbiased reasoning while preserving general capability
- Fine-tune models for verifiable domains such as mathematics, code, games, structured prediction, or other programmatically evaluated tasks
Research Evaluation & Optimisation
- Design statistically sound experiments and benchmark comparisons
- Evaluate training efficiency, scaling behaviour, and generalisation
- Develop contamination controls and robust model-evaluation protocols
- Analyse model failures and propose targeted training or data interventions
- Document research findings, experimental methodology, and technical conclusions clearly
Ideal Profile Strong candidates may have:
- At least 3 years of machine learning research experience, including qualifying doctoral research
- Hands-on experience training or fine-tuning transformer-based language models
- Strong expertise in one or more of foundation model pre-training, pre-training data, or LLM post-training
- Experience working with PyTorch, JAX, TensorFlow, or comparable machine learning frameworks
- Ability to design and execute empirical research independently
- Strong understanding of optimisation, evaluation methodology, and experimental design
- Excellent technical writing, analytical reasoning, and research communication skills
- Experience working with large-scale datasets and distributed training systems
Educational Background
- A degree in computer science, machine learning, artificial intelligence, mathematics, statistics, engineering, or a related discipline is highly relevant
- PhD research in machine learning, natural language processing, deep learning, or a related field may count towards the experience requirement
- A strong publication record, impactful open-source contributions, or comparable applied research experience may also be considered
- Research experience at a leading university, technology company, AI organisation, or research laboratory may strengthen an application
Nice to Have
- Research experience involving scaling laws or training efficiency
- Familiarity with curriculum learning, data ordering, and mixture optimisation
- Experience constructing LLM benchmarks and controlling for training-data contamination
- Background in reinforcement learning for language models
- Expertise in reward modelling, preference learning, or human-feedback pipelines
- Experience with model alignment, AI safety, truthfulness, or refusal behaviour
- Familiarity with synthetic data generation and weak-supervision methods
- Publications or significant open-source contributions related to foundation models or language-model training
Why This Opportunity
- Work on cutting-edge foundation model research
- Investigate challenging empirical problems across pre-training and post-training
- Apply advanced machine learning expertise to high-impact language-model development
- Collaborate asynchronously with experienced AI researchers
- Explore methods for improving model capability, efficiency, reliability, and alignment
- Participate in flexible project-based work with competitive hourly compensation
Contract Details
- Independent contractor role
- Fully remote with flexible scheduling
- Competitive rates between $95-$115 per hour depending on expertise and project scope
- Work may include model training, dataset development, post-training pipeline design, evaluation, and experimental research
- Weekly payments via Stripe or Wise
- Projects may be extended, shortened, or adjusted depending on scope and performance
- Work will not involve access to confidential or proprietary information from any employer, client, or institution
About the Platform This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.