Senior ML Engineer
$102K - $141K/yr
Evaluate, benchmark, and select inference providers (e.g., Together AI, Fireworks, Groq, Replicate, AWS Bedrock, Azure OpenAI) based on latency, cost, throughput, and model capability trade-offs.
Quick apply
$102K - $141K/yr
Evaluate, benchmark, and select inference providers (e.g., Together AI, Fireworks, Groq, Replicate, AWS Bedrock, Azure OpenAI) based on latency, cost, throughput, and model capability trade-offs.
Quick apply
$102K - $141K/yr
Evaluate, benchmark, and select inference providers (e.g., Together AI, Fireworks, Groq, Replicate, AWS Bedrock, Azure OpenAI) based on latency, cost, throughput, and model capability trade-offs.
Tampa, FL · On-site
Architect and implement scalable ML training and inference pipelines using AWS SageMaker, managing model training, hyperparameter tuning, distributed training for large vision models, and real-time ...
Tampa, FL · On-site
Architect and implement scalable ML training and inference pipelines using AWS SageMaker, managing model training, hyperparameter tuning, distributed training for large vision models, and real-time ...
Architect and implement scalable ML training and inference pipelines using AWS SageMaker, managing model training, hyperparameter tuning, distributed training for large vision models, and real-time ...
Architect and implement scalable ML training and inference pipelines using AWS SageMaker, managing model training, hyperparameter tuning, distributed training for large vision models, and real-time ...
Architect and implement scalable ML training and inference pipelines using AWS SageMaker, managing model training, hyperparameter tuning, distributed training for large vision models, and real-time ...
Architect and implement scalable ML training and inference pipelines using AWS SageMaker, managing model training, hyperparameter tuning, distributed training for large vision models, and real-time ...
Tampa, FL · On-site
Architect and implement scalable ML training and inference pipelines using AWS SageMaker, managing model training, hyperparameter tuning, distributed training for large vision models, and real-time ...
Tampa, FL · On-site
Architect and implement scalable ML training and inference pipelines using AWS SageMaker, managing model training, hyperparameter tuning, distributed training for large vision models, and real-time ...
Tampa, FL · On-site
Architect and implement scalable ML training and inference pipelines using AWS SageMaker, managing model training, hyperparameter tuning, distributed training for large vision models, and real-time ...
Tampa, FL · On-site
Architect and implement scalable ML training and inference pipelines using AWS SageMaker, managing model training, hyperparameter tuning, distributed training for large vision models, and real-time ...
Tampa, FL · On-site
... inference, and monitoring in production environments. • Participate in continuous improvement of the ML infrastructure and processes for scalability and performance. Qualifications : Required : • ...
Tampa, FL · On-site
... inference, and monitoring in production environments. • Participate in continuous improvement of the ML infrastructure and processes for scalability and performance. Qualifications : Required : • ...
Jacksonville, FL · On-site
... inference, and monitoring in production environments. • Participate in continuous improvement of the ML infrastructure and processes for scalability and performance. Qualifications : Required : • ...
Jacksonville, FL · On-site
... inference, and monitoring in production environments. • Participate in continuous improvement of the ML infrastructure and processes for scalability and performance. Qualifications : Required : • ...
Miami, FL · On-site
... inference, and monitoring in production environments. • Participate in continuous improvement of the ML infrastructure and processes for scalability and performance. Qualifications : Required : • ...
Miami, FL · On-site
... inference, and monitoring in production environments. • Participate in continuous improvement of the ML infrastructure and processes for scalability and performance. Qualifications : Required : • ...
Develop causal inference methodologies to understand true incrementality of product changes ... Proven track record building and deploying ML models in production , particularly in ...
Quick apply
Develop causal inference methodologies to understand true incrementality of product changes ... Proven track record building and deploying ML models in production , particularly in ...
Job : Backend Python Developer - AI/ML Location : Orlando preferable, Las Vegas Skills : Python ... Experience integrating LLM APIs (OpenAI, HuggingFace Inference API, etc.) into real-world ...
Quick apply
Job : Backend Python Developer - AI/ML Location : Orlando preferable, Las Vegas Skills : Python ... Experience integrating LLM APIs (OpenAI, HuggingFace Inference API, etc.) into real-world ...
Tampa, FL · On-site
... inference operations. Qualifications : Required : • 3+ years of relevant hands-on experience • ... and maintaining ML pipelines • Practical experience with PyTorch (TensorFlow experience ...
Tampa, FL · On-site
... inference operations. Qualifications : Required : • 3+ years of relevant hands-on experience • ... and maintaining ML pipelines • Practical experience with PyTorch (TensorFlow experience ...
$113K - $149K/yr
Integrate AI/ML capabilities into production systems (e.g., model inference APIs, decision-support features, anomaly detection workflows) * Design and optimize data models and persistence layers to ...
Quick apply
$113K - $149K/yr
Integrate AI/ML capabilities into production systems (e.g., model inference APIs, decision-support features, anomaly detection workflows) * Design and optimize data models and persistence layers to ...
Review and assess new ML/statistical techniques for practical trading relevance What We're Looking ... Deep knowledge of statistical inference, hypothesis testing, and overfitting risks * Experience ...
Review and assess new ML/statistical techniques for practical trading relevance What We're Looking ... Deep knowledge of statistical inference, hypothesis testing, and overfitting risks * Experience ...
... inference systems and low-latency model serving Knowledge of adversarial ML and AI security/robustness techniques Experience with graph neural networks for network analysis Experience in design ...
... inference systems and low-latency model serving Knowledge of adversarial ML and AI security/robustness techniques Experience with graph neural networks for network analysis Experience in design ...
Orlando, FL · On-site
$95K - $126K/yr
Oversee enterprisescale AI platforms supporting model training, inference, evaluation, monitoring ... Leadershiplevel expertise in AI/ML platform engineering, spanning MLOps, LLMOps, and AIOps.
Orlando, FL · On-site
$95K - $126K/yr
Oversee enterprisescale AI platforms supporting model training, inference, evaluation, monitoring ... Leadershiplevel expertise in AI/ML platform engineering, spanning MLOps, LLMOps, and AIOps.
Naples, FL · On-site
$96K - $127K/yr
Oversee enterprisescale AI platforms supporting model training, inference, evaluation, monitoring ... Leadershiplevel expertise in AI/ML platform engineering, spanning MLOps, LLMOps, and AIOps.
Naples, FL · On-site
$96K - $127K/yr
Oversee enterprisescale AI platforms supporting model training, inference, evaluation, monitoring ... Leadershiplevel expertise in AI/ML platform engineering, spanning MLOps, LLMOps, and AIOps.
Jacksonville, FL · On-site
$106K - $127K/yr
Senior AIOps ML Engineer Descriptions: "Core Responsibilities Lakehouse Architecture & Data ... Build low-latency streaming inference pipelines (Flink / Spark Streaming) for real-time anomaly ...
Jacksonville, FL · On-site
$106K - $127K/yr
Senior AIOps ML Engineer Descriptions: "Core Responsibilities Lakehouse Architecture & Data ... Build low-latency streaming inference pipelines (Flink / Spark Streaming) for real-time anomaly ...
Jacksonville, FL · On-site
$100K - $136K/yr
Experience supporting GPU-based workloads and optimizing infrastructure for AI model training and inference.This role will be instrumental in establishing and evolving the organization's AI/ML ...
Jacksonville, FL · On-site
$100K - $136K/yr
Experience supporting GPU-based workloads and optimizing infrastructure for AI model training and inference.This role will be instrumental in establishing and evolving the organization's AI/ML ...
Fort Lauderdale, FL · On-site
Apply causal inference methods to understand the impact of potential product changes. * Define and build new ML features using text and multimodal embeddings and GenAI. * Validate offline learnings ...
Fort Lauderdale, FL · On-site
Apply causal inference methods to understand the impact of potential product changes. * Define and build new ML features using text and multimodal embeddings and GenAI. * Validate offline learnings ...
| Aspect | ML Inference | Data Scientist |
|---|---|---|
| Required Credentials | Knowledge of machine learning models, programming skills | Degree in data science, statistics, or related fields |
| Work Environment | Deploying models in production, real-time data processing | Data analysis, model development, research |
| Industry Usage | AI product deployment, software companies | Research institutions, tech firms, consulting |
ML Inference focuses on deploying trained models to make predictions on new data, often in real-time. Data Scientists develop and analyze models, working primarily in research and development. While both roles require understanding of machine learning, ML Inference emphasizes deployment and operationalization, whereas Data Scientists focus on model creation and analysis.
For Ml Inference jobs in Florida, the most frequently searched job titles are:
The top searched job categories for Ml Inference jobs in Florida are:
Cities in Florida with the most Ml Inference job openings:
$102K - $141K/yr
Full-time
Medical, Dental, Vision, Life, Retirement, PTO
Posted 18 days ago
About IntelePeer.ai:
IntelePeer is a healthcare-focused AI communications platform that powers AI voice agents and intelligent workflow automation for ambulatory care groups, all specialty healthcare verticals, health systems, and payers. Our AI Agent suite and SmartFlow platform are deployed at scale across some of the nation's most complex healthcare organizations — handling millions of patient interactions annually for scheduling, care coordination, billing inquiry, and more. We build AI that talks to real patients and produces real outcomes, and we need people who take that responsibility seriously.
Job Summary:
IntelePeer is building AI-native communications products and we need an ML engineer who gets their hands dirty. This is not a research role — you will own the full lifecycle of machine learning systems: designing training pipelines, fine-tuning and aligning large language models, optimizing inference, and shipping models that run reliably in production. You will work alongside our AI Engineering team to push the capabilities of our platform and deliver measurable impact.
Responsibilities:
• Design, implement, and maintain end-to-end ML training pipelines — from raw data ingestion and preprocessing through model training, evaluation, and deployment.
• Fine-tune large language models using techniques such as LoRA, QLoRA, and full fine-tuning; apply PEFT strategies to balance performance and compute cost.
• Implement and experiment with reinforcement learning from human feedback (RLHF) workflows, including PPO (Proximal Policy Optimization) and GRPO (Group Relative Policy Optimization) for model alignment and preference optimization.
• Host, serve, and optimize LLMs in production using inference frameworks such as vLLM, Text Generation Inference (TGI), Triton Inference Server, or ONNX Runtime.
• Evaluate, benchmark, and select inference providers (e.g., Together AI, Fireworks, Groq, Replicate, AWS Bedrock, Azure OpenAI) based on latency, cost, throughput, and model capability trade-offs.
• Build and maintain embedding pipelines — generate, index, and retrieve dense embeddings using vector databases (Pinecone, pgvector, Weaviate, or similar) for RAG and semantic search applications.
• Implement and expose ML capabilities via Model Context Protocol (MCP) — enabling AI agents to call model-backed tools in a structured, context-aware manner.
• Perform rigorous data analysis and processing: clean, transform, and curate datasets for training, fine-tuning, and evaluation; build data quality and validation pipelines.
• Develop robust model evaluation frameworks — define metrics, build eval harnesses, run A/B experiments, and track regressions across model versions.
• Collaborate with software engineers to integrate ML systems into product features via FastAPI services; ensure models are observable, versioned, and maintainable in production.
Supervisory Duties: This is an IC role
Minimum Education and Experience:
Bachelors in computer science or statistics
• 3–8+ years of hands-on ML engineering experience with a strong production track record.
• Deep understanding of core ML concepts: neural network architectures (transformers, attention mechanisms), loss functions, optimization algorithms, regularization, and model evaluation.
• Practical experience fine-tuning LLMs (LoRA, QLoRA, PEFT, instruction tuning, DPO) on custom datasets using frameworks such as Hugging Face Transformers, TRL, or Axolotl.
• Hands-on experience with RL-based alignment techniques — specifically PPO and GRPO — for reward modeling, preference optimization, and RLHF pipelines.
• Experience hosting and serving LLMs: vLLM, TGI, Triton, or similar; understanding of model quantization (GPTQ, AWQ, int4/int8), batching strategies, and throughput optimization.
• Working knowledge of major inference vendors and cloud AI APIs; ability to evaluate and select providers based on cost, latency, and capability benchmarks.
• Proficiency in embedding models (sentence-transformers, OpenAI embeddings, or equivalent) and vector search infrastructure for RAG pipelines.
• Understanding of Model Context Protocol (MCP) and how to expose ML functionality as structured tools for agentic systems.
Key Competencies:
• Experience with distributed training frameworks (DeepSpeed, FSDP, Megatron-LM) for multi-GPU or multi-node training runs.
• Familiarity with MLOps tooling: MLflow, Weights & Biases, DVC, or similar for experiment tracking, model registry, and pipeline orchestration.
• Knowledge of synthetic data generation techniques for augmenting fine-tuning datasets.
• Exposure to multimodal models (vision-language, speech-language) or voice/speech AI systems.
• Contributions to open-source ML projects or published research (papers, blog posts, or technical write-ups).
Physical Requirements:
· Sedentary work lifting no more than 10 pounds.
· Occasional lifting, carrying, and standing.
· Frequent hand/eye coordination to operate office equipment.
· Vision sufficient to read computer screens, reports, and related department documents.
· Dexterity to operate computer keyboards and other related office equipment.
· Endurance sufficient to sit and work at a computer for extended periods of time.
· Frequent speech communication and hearing.
Why you'll love it here:
Unlimited Vacation for exempt employees
Paid Holidays
Competitive medical, dental & vision insurance for employees and their dependents
401K Retirement Plan
Stock Options
Company-paid life insurance
Health & Flexible Savings Accounts
Cell phone, gym, and internet reimbursement
Paid Parental Leave
Tuition Reimbursement
Employee Assistance Program (EAP)
Free snacks (Denver, and or Fort Lauderdale)
Fun events (virtual and in-person)
Applicants must be authorized to work for any employer in the U.S.
We are unable to sponsor or take over sponsorship of an employment visa at this time.
Any requests to exercise your rights as a data subject under GDPR should be submitted to infosec@intelepeer.com for prompt processing. Please refer to our Privacy Policy (at www.intelepeer.com/privacy/intelepeer-privacy-policy) for any questions on how IntelePeer complies with GDPR.
For California residents only: Please refer to the link below for IntelePeer’s Applicant CCPA Privacy Notice. https://intelepeer.com/privacy/intelepeer-california-applicant-privacy-notice/
IntelePeer participates in E-Verify.
https://www.eeoc.gov/poster
At IntelePeer, we value diversity and are proud to be an Equal Opportunity Employer. We do not discriminate on the basis of race, color, religion, sex, national origin, age, disability, genetic information, or any other protected status.
We strive to provide reasonable accommodations to applicants and employees with disabilities to support them in performing the essential functions of their roles.
If you have any questions or need assistance, please contact our Director of Recruiting.
Sourced by ZipRecruiter
Software development
201 - 500 Employees
San Mateo, CA, US