1

Manager Tensor Jobs in New York (NOW HIRING)

Implementing parallelization strategies (data, tensor, pipeline, context) and optimizing ... Deep understanding of GPU memory management and distributed systems profiling * Hands-on RL ...

Senior AWS Python Developer

New York, NY ยท On-site

$132K - $178K/yr

... Manage individual project priorities deadlines and deliverables Ensure that solutions are in line ... PyTorch Tensor Flow is strongly preferred Familiarity with the agile process stand ups planning ...

next page

Showing results 1-20

Manager Tensor information

What is the difference between Manager Tensor vs Data Scientist?

AspectManager TensorData Scientist
Required CredentialsBachelor's or Master's in Computer Science, Data Analytics, or related fields; certifications like TensorFlow Developer are commonBachelor's or Master's in Data Science, Statistics, Computer Science; certifications like Certified Data Scientist are common
Work EnvironmentLeads teams, manages projects, collaborates with stakeholders in tech or AI-focused companiesAnalyzes data, builds models, reports insights in tech, finance, healthcare industries
Employer & Industry UsageUsed in AI, machine learning, and tech companies for managing TensorFlow projectsUsed across industries for data analysis, predictive modeling, and research

The main difference is that a Manager Tensor oversees AI projects involving TensorFlow, focusing on team management and project delivery, while a Data Scientist primarily analyzes data and builds models. Both roles require technical knowledge, but the Manager Tensor role emphasizes leadership and project management within AI initiatives.

What are the key skills and qualifications needed to thrive as a Manager Tensor, and why are they important?

To thrive as a Manager Tensor (commonly referred to as a TensorFlow Manager or Machine Learning Manager), you need a solid background in machine learning, deep learning frameworks (especially TensorFlow), and experience leading technical teams, typically backed by a relevant degree. Proficiency with TensorFlow, Python, data engineering tools, and cloud platforms, along with certifications in machine learning, are highly valued. Leadership, strong communication, and project management skills help you effectively guide teams and collaborate with stakeholders. These skills ensure successful project delivery, innovation, and alignment with organizational goals in complex AI-driven environments.

What are some common challenges faced by a Manager Tensor when leading AI and machine learning teams?

A Manager Tensor often encounters challenges such as balancing technical leadership with strategic oversight, managing projects that involve complex and evolving technologies, and ensuring effective communication among data scientists, engineers, and stakeholders. Additionally, staying current with rapid advancements in AI frameworks and guiding the team through best practices can be demanding. Collaboration across multidisciplinary teams and aligning projects with business objectives are also key aspects of the role.

What is a Manager Tensor?

A Manager Tensor is typically a managerial position responsible for overseeing teams that develop and implement machine learning models using TensorFlow or similar tensor-based frameworks. This role involves coordinating data science and engineering teams, ensuring project goals align with business objectives, and facilitating the deployment of scalable AI solutions. Additionally, a Manager Tensor may be tasked with mentoring staff, managing resources, and staying updated with the latest advancements in artificial intelligence. The position requires strong leadership, technical expertise in machine learning, and experience with deep learning platforms.
What cities in New York are hiring for Manager Tensor jobs? Cities in New York with the most Manager Tensor job openings:

Member of Technical Staff (AI Inference Engineer)

Perplexity

New York, NY โ€ข On-site

$220K - $485K/yr

Full-time

Re-posted 12 days ago


Job description

We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL - and we need another engineer to join us.
What you will work on
Examples of real work the team does:
  • New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway.
  • GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow.
  • Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic.
  • Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernel interleaving.
  • Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents.

Who we're looking for
  • Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar). Any other deep systems programming experience is a plus.
  • You understand modern LLM architectures and are able to bring them up reliably in a production environment.
  • You've built and operated production distributed systems under real load - ideally performance-critical ones.
  • Comfortable working across languages and layers: Rust for the serving runtime, Python for model code, CUDA/CuteDSL for kernels.
  • You own problems end-to-end. You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on Friday.
  • Self-directed. You do well in fast-moving environments where the path forward isn't laid out for you.

Good if you touched any of
  • ML compilers and framework internals: PyTorch internals, torch.compile, custom operators.
  • Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism.
  • Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving.
  • Profiling and debugging tools: Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis.
  • Container orchestration: Kubernetes, GPU scheduling, autoscaling inference workloads.

Qualifications
  • 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
  • Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).
  • Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores).
  • Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation).