1

Deepspeed Jobs in Florida (NOW HIRING)

Machine Learning Engineer

Orlando, FL · On-site

$120 - $160/hr

DeepSpeed, Accelerate, Ray, distributed training frameworks Models: GPT/LLaMA variants, DALL‑E/Stable Diffusion, Whisper, multi‑modal models Fine‑tuning: LoRA, QLoRA, DreamBooth, custom ...

Senior ML Engineer

Dania Beach, FL · On-site

$102K - $141K/yr

Experience with distributed training frameworks (DeepSpeed, FSDP, Megatron-LM) for multi-GPU or multi-node training runs. * Familiarity with MLOps tooling: MLflow, Weights & Biases, DVC, or similar ...

Deepspeed information

What are some common challenges faced by engineers working with DeepSpeed and how can they be addressed?

Engineers working with DeepSpeed often encounter challenges related to optimizing large-scale model training, such as managing memory efficiency and tuning distributed training parameters. Troubleshooting issues like gradient accumulation, parallelism strategies, and ensuring compatibility with different hardware setups can be complex. Collaborating closely with data scientists, DevOps, and research teams is essential for addressing these challenges, as is staying updated with the latest DeepSpeed releases and documentation. Regular participation in code reviews and knowledge-sharing sessions can also help engineers overcome technical hurdles and continuously improve model performance.

What is DeepSpeed?

Deepspeed is an open-source deep learning optimization library developed by Microsoft, designed to enable distributed training of large-scale models efficiently. It helps researchers and engineers train models that are too large to fit in the memory of a single GPU by offering features like ZeRO optimization, mixed-precision training, and advanced parallelism techniques. Deepspeed is widely used in the machine learning community for its scalability and performance improvements, making it easier to train state-of-the-art models on vast datasets. The library integrates seamlessly with PyTorch and supports training on multiple GPUs and even across multiple machines.

What is the difference between Deepspeed vs Data Scientist?

AspectDeepspeedData Scientist
Required credentialsKnowledge of machine learning frameworks, programming skills in Python, experience with AI model trainingDegree in Data Science, Statistics, Computer Science, or related fields; strong analytical skills
Work environmentAI research labs, tech companies, cloud computing environmentsBusiness, tech companies, research institutions
Industry usageAI model training, deep learning optimizationData analysis, predictive modeling, business insights

Deepspeed focuses on optimizing large-scale AI model training and deep learning performance, while Data Scientists analyze data to generate insights and build predictive models. Both roles require technical skills but serve different purposes within the AI and data ecosystem.

What are the key skills and qualifications needed to thrive as a DeepSpeed engineer, and why are they important?

To thrive as a DeepSpeed Engineer, you need a solid background in machine learning, deep learning frameworks (such as PyTorch), and distributed systems, often supported by a degree in computer science or a related field. Proficiency with DeepSpeed, parallel computing libraries, and cloud platforms, along with familiarity with tools like CUDA and NCCL, is typically expected. Strong problem-solving abilities, collaboration, and adaptability are crucial soft skills for optimizing large-scale AI models and working with cross-functional teams. Mastering these skills ensures efficient development and deployment of high-performance, scalable AI solutions in demanding environments.
What job categories do people searching Deepspeed jobs in Florida look for? The top searched job categories for Deepspeed jobs in Florida are:
What cities in Florida are hiring for Deepspeed jobs? Cities in Florida with the most Deepspeed job openings:

Machine Learning Engineer

247Hire

Orlando, FL • On-site

$120 - $160/hr

Other

Posted 5 days ago


Job description

Seeking a Machine Learning Engineer for the following role - Generative AI & ML Frameworks: PyTorch, TensorFlow, Hugging Face Transformers, Diffusers Training: DeepSpeed, Accelerate, Ray, distributed training frameworks Models: GPT/LLaMA variants, DALL‑E/Stable Diffusion, Whisper, multi‑modal models Fine‑tuning: LoRA, QLoRA, DreamBooth, custom training pipelines Infrastructure & Platforms Cloud: GCP Vertex AI, Azure OpenAI, AWS Bedrock, multi‑cloud orchestration Serving: TensorRT, ONNX, TorchServe, custom inference servers Orchestration: Kubernetes, Docker, APIGEE, Terraform Data: Vector databases (Pinecone, Weaviate), feature stores, data versioning Specialized Tools Frameworks: Autogen, LangChain, MCP (Model Context Protocol) Evaluation: Custom metrics, human evaluation platforms, A/B testing frameworks Monitoring: MLflow, Weights & Biases, custom dashboards

Responsibilities
  • Build text‑to‑image and text‑to‑video generation systems
  • Develop speech synthesis and voice cloning models with safety guardrails for character voices
  • Create image‑to‑text and video‑to‑text systems for content analysis and accessibility
  • Implement cross‑modal generation (text + image? video, audio + text? multimedia content)
  • Build real‑time generative systems for interactive experiences (IoT)
  • Model Evaluation & Quality Assurance
  • Design and implement custom evaluation models for content assessment (brand safety, content ratings, character consistency)
  • Build automated benchmarking systems for generative model performance across multi‑cloud environments
  • Develop specialized ML pipelines for hallucination detection, bias measurement, and factual accuracy assessment
  • Create domain‑specific evaluation frameworks for use cases (content appropriateness, brand alignment, safety compliance)
  • Implement human‑in‑the‑loop evaluation systems with domain experts
  • Research & Advanced Techniques: Implement cutting‑edge generative AI techniques: diffusion models, transformer variants, mixture of experts
  • Develop constitutional AI and AI safety techniques for responsible content generation
  • Build adversarial training systems to improve model robustness
  • Research and implement prompt engineering and in‑context learning optimization
  • Create novel architectures for specific generative tasks
  • Production AI/ML Systems: Design A/B testing frameworks for generative model comparison and optimization
  • Build real‑time inference optimization for low‑latency content generation
  • Implement model serving infrastructure with auto‑scaling and load balancing
  • Create model monitoring, drift detection, and automatic retraining systems
  • Develop caching and retrieval systems for improved generative AI performance
Key Projects & Use Cases (Marketing Content Generation)
  • Build text‑to‑video systems for promotional content creation
  • Develop brand‑consistent image generation with style transfer
  • Create voice synthesis for character‑based marketing campaigns
Theme Park Innovation
  • Implement real‑time generative systems for interactive guest experiences
  • Build personalized content generation based on guest preferences
  • Develop safety‑aware content generation for operational communications
Customer Experience Enhancement
  • Create personalized response generation for customer support
  • Build multi‑lingual content generation for global audiences
  • Develop accessibility‑focused content generation (audio descriptions, simplified language)
Basic Qualifications
  • 5+ years of hands‑on machine learning engineering with 2+ years focused on generative AI
  • Strong experience with transformer architectures, diffusion models, and large language models
  • Proven track record with model fine‑tuning, RLHF, and parameter‑efficient training techniques
  • Experience with multi‑modal AI systems (text+vision, text+audio, cross‑modal generation)
  • Deep understanding of generative AI training dynamics, loss functions, and optimization techniques
Technical Expertise
  • Expert‑level Python programming with TensorFlow/PyTorch and distributed training frameworks
  • Experience with cloud ML platforms (GCP Vertex AI, Azure OpenAI, AWS Bedrock) and model serving
  • Strong background in computer vision, NLP, and audio processing for generative applications
  • Knowledge of MLOps, model versioning, and production deployment strategies
  • Experience with vector databases, embeddings, and retrieval‑augmented generation (RAG)
AI Safety & Evaluation
  • Experience building evaluation frameworks for generative AI systems
  • Knowledge of AI safety techniques: bias detection, content filtering, adversarial robustness
  • Understanding of responsible AI frameworks and red‑team methodologies
  • Familiarity with AI governance, model interpretability, and compliance requirements
Preferred Qualifications
  • Advanced degree in Machine Learning, Computer Science, or related field
  • Experience with industry applications (content creation, media analysis, interactive systems)
  • Knowledge of edge AI optimization and real‑time inference systems
  • Background in reinforcement learning and human preference modeling
  • Experience with large‑scale distributed training (multi‑GPU, multi‑node)
  • Contributions to open‑source AI projects or published research in generative AI
Education

BE/BS in Machine Learning, Computer Science, or related field

#J-18808-Ljbffr

247Hire logo

About 247Hire

Sourced by ZipRecruiter

Industry

Recruiting and staffing services

Company size

201 - 500 Employees

Headquarters location

Oak Brook, IL, US

Year founded

2002