1

Multimodal Learning Jobs in Conroe, TX (NOW HIRING)

Design and implement deep-learning solutions using neural networks, transformers, convolutional architectures, sequence models, representation-learning techniques, and multimodal approaches. * Build ...

Senior Software Engineer

Spring, TX · On-site

$109K - $143K/yr

Audio-video multimodal models * Human computer intelligent interactions * Deep neural networks ... Integrate, evaluate, and deploy machine learning models - including LLMs, vision models, and audio ...

Senior Software Engineer

Spring, TX

$109K - $143K/yr

Audio-video multimodal models * Human computer intelligent interactions * Deep neural networks ... Integrate, evaluate, and deploy machine learning models - including LLMs, vision models, and audio ...

next page

Showing results 1-20

Multimodal Learning information

See Conroe, TX salary details

$18K

$52.8K

$98K

How much do multimodal learning jobs pay per year?

As of Sep 14, 2026, the average yearly pay for multimodal learning in Conroe, TX is $52,816.00, according to ZipRecruiter salary data. Most workers in this role earn between $35,100.00 and $61,600.00 per year, depending on experience, location, and employer.

What is multimodal learning?

Multimodal learning is an area of machine learning that involves integrating and processing information from multiple types of data, such as text, images, audio, and video. The goal is to create models that can understand and make predictions based on more than one data modality, similar to how humans use various senses. This approach is used in applications like speech recognition with visual cues, image captioning, and video analysis. By combining different data types, multimodal learning systems can achieve better accuracy and more robust understanding.

What are the key skills and qualifications needed to thrive in multimodal learning, and why are they important?

To excel as a Multimodal Learning Specialist, you need a solid background in machine learning, data science, and computer vision, often supported by an advanced degree in a related field. Familiarity with deep learning frameworks like TensorFlow or PyTorch, experience integrating data from diverse sources (e.g., text, audio, images), and knowledge of relevant algorithms are crucial. Strong problem-solving abilities, creativity, and effective collaboration are standout soft skills for this role. These competencies are vital for developing innovative models that can process and interpret complex, multi-source data to drive impactful AI solutions.

What are some common challenges faced by professionals working in multimodal learning roles, and how can they be addressed?

Professionals in multimodal learning frequently encounter challenges related to integrating and aligning data from multiple sources, such as text, images, audio, or video. Ensuring data quality and consistency across modalities can be complex, and developing models that effectively combine heterogeneous information often requires advanced technical skills and innovative thinking. Collaboration with domain experts and other data scientists is key to overcoming these obstacles, as is staying up to date with the latest research and tools in machine learning. Regular team meetings and cross-disciplinary workshops can help foster a collaborative environment and promote knowledge sharing.

What is the difference between Multimodal Learning vs Data Scientist?

AspectMultimodal LearningData Scientist
Required CredentialsAdvanced degrees in AI, Machine Learning, or Computer ScienceBachelor's or Master's in Data Science, Statistics, or related fields
Work EnvironmentResearch labs, AI development teams, academiaBusiness, tech companies, analytics teams
Industry UsageAI research, multimedia applications, roboticsData analysis, predictive modeling, business insights

Multimodal Learning focuses on developing AI models that process and integrate multiple data types like images, text, and audio. Data Scientists analyze data to extract insights, build models, and support decision-making. While both roles involve data and algorithms, Multimodal Learning is specialized in AI model development for complex data integration, whereas Data Scientists work broadly across data analysis and interpretation.

What job categories do people searching Multimodal Learning jobs in Conroe, TX look for?

The top searched job categories for Multimodal Learning jobs in Conroe, TX are:

What cities near Conroe, TX are hiring for Multimodal Learning jobs?

Cities near Conroe, TX with the most Multimodal Learning job openings:

Robotics Data Pipeline Engineer - Multimodal Data

Houston, TX • On-site

$109K - $131K/yr

Other

Medical, PTO

Re-posted yesterday


Job description

Job Title: Robotics Data Pipeline Engineer – Multimodal Data

Department: Software

Reports To: Teleoperations Lead

Employment Type: Full-Time

Location: Houston, TX or Pensacola Fl

Who We Are

Persona AI is building humanoid robots for the most demanding environments in heavy industry — shipyards, steel mills, fabrication facilities, and offshore platforms — performing welding, grinding, maintenance, inspection, and material‑handling work that is dangerous, physically demanding, and increasingly difficult to staff.

We are backed by leading strategic and financial investors and engaged with global industrial leaders across Korea, Japan, the United States, and Singapore. Korea is the center of gravity for our early commercial strategy, anchored by relationships with the world’s leading shipbuilders and steelmakers. Our work spans both the robot platform itself and the systems, partners, and playbooks required to deploy it at scale.

Why Join Persona AI?
  • We offer competitive compensation, a performance-based bonus, 99% employer covered medical benefits, early‑stage equity, competitive PTO, and a company‑wide paid winter break between December 24th and January 2nd.
  • You’ll shape technology that’s redefining the possibilities of robotics and human interaction.
  • Work alongside passionate teammates who value creativity, and continuous learning.
  • Enjoy full access to advanced tools,
About the Role

As a Data Pipeline Engineer, you will architect and scale the data infrastructure that feeds our foundation models. Your primary mission is to extract, augment, and align human dexterous manipulation data from massive complex, multi‑sensor and egocentric video datasets. Crucially, you will build advanced post‑processing algorithms to perform deep force analysis and infer hidden states from raw data—such as processing direct force‑torque outputs to quantify grasp dynamics, estimating contact forces from visual cues, extrapolating heavily occluded hand positions, or deriving 3D geometry from 2D frames. You will use spatial, temporal, and cross‑modal data augmentation to multiply the value of every minute of data our teleoperation team collects.

What You Will Be Doing
  • Multimodal Data Pipelines: Architect end‑to‑end ingestion pipelines that take raw, unstructured recordings—egocentric video, teleoperation sessions, third‑party open datasets—and produce indexed, queryable, training‑ready datasets. This includes temporal segmentation of long recordings into action clips, metadata and scene‑graph extraction, embedding‑based retrieval, and language annotation workflows.
  • Force Analysis & Hidden State Inference: Design cross‑modal validation systems that verify video, proprioception, force/haptic signals, and language annotations agree with each other—e.g., reprojecting robot state into the image plane to confirm video–state consistency, and VLM‑assisted checks that instructions match observed behavior.
  • Kinematic Retargeting & Alignment: Orchestrating hand‑tracking, segmentation, depth estimation, 3D reconstruction, and pose‑tracking modules; retargeting human demonstrations into robot trajectories; and running simulation‑in‑the‑loop validation (kinematic feasibility, physics replay, motion‑consistency filtering) so synthesized data is physically grounded, not just visually plausible.
  • Advanced Data Augmentation: Implement robust data augmentation strategies (spatial transformations, temporal scaling, synthetic viewpoints, and sensor noise injection) to expand expert trajectories and improve the robustness of our learning models.
  • Teleoperation Synchronization: Unified state–action representations across differing embodiments, coordinate frames, rotation conventions, gripper/hand parameterizations, and sampling rates—with per‑dimension validity masking and per‑source normalization so that adding a new robot or sensor is a configuration change, not a rewrite.
  • Close the loop with data consumers: Build the tooling that lets researchers query, visualize, and audit datasets (clip browsers, trajectory viewers, annotation review UIs), and turn model‑failure analyses into new curation rules and targeted re‑collection requests.
What We Are Looking For
  • Education: M.S., or Ph.D. in Computer Science, Data Engineering, Machine Learning, Robotics, Mechanical Engineering, or a related field.
  • Programming & ML Frameworks: Deep expertise in Python and extensive experience with PyTorch, specifically in handling custom dataloaders for multimodal datasets.
  • Force & Time‑Series Data Processing: Experience analyzing and processing complex time‑series data from force‑torque (F/T) sensors, load cells, or tactile arrays, ensuring pristine alignment with visual frames.
  • Video Processing Expertise: Mastery of video processing pipelines and libraries (OpenCV, FFmpeg, Decord) and managing the I/O bottlenecks of terabyte‑scale video datasets.
  • Solid working knowledge of 3D geometry and robotics data: coordinate frames and transforms, rotation representations, camera intrinsics/extrinsics, forward/inverse kinematics, URDF—enough to build automated checks that catch geometric inconsistencies in the data.
  • Data Augmentation: Proven ability to implement programmatic and generative data augmentation techniques for computer vision and time‑series data.
Bonus Skills
  • Experience with NVIDIA’s robotic software stack (Open X-Embodiment, DROID, AgiBot World, EgoDex, or similar).
  • Familiarity with the modern perception toolbox as a user: segmentation (SAM‑family), monocular depth, hand/body pose estimation (MANO/SMPL), 6‑DoF object pose tracking, point tracking—you don't need to train these models, but you should be comfortable composing and evaluating them in a pipeline.
  • Familiarity with distributed data processing systems (Ray, Apache Spark) for cluster computing.
  • Background in generating or utilizing synthetic robotic data via simulation (Omniverse, MuJoCo).
  • Experience integrating spatial awareness or tactile data representations (e.g., Fourier encoding) into visual pipelines.

Persona AI is an Equal Opportunity Employer.

All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, age, disability, veteran status, or any other characteristic protected by applicable federal, state, or local law.

#J-18808-Ljbffr