1

Internship Scale Ai Jobs (Flexible Options) Near Me

... scale problems. Your Impact As a Senior AI Researcher you will operate as a visionary technical ... Mentor researchers, engineers and interns, elevating the team's scientific rigor and engineering ...

... scale problems. Your Impact As a Senior AI Researcher you will operate as a visionary technical ... Mentor researchers, engineers and interns, elevating the team's scientific rigor and engineering ...

Applied Data Science Summer Internship About Us: Evolver is a rapidly growing enterprise AI company ... Access enterprise-scale AI and data science initiatives through hands-on collaboration with ...

Applied Data Science Summer Internship About Us: Evolver is a rapidly growing enterprise AI company ... Access enterprise-scale AI and data science initiatives through hands-on collaboration with ...

ALAB) provides rack-scale AI infrastructure through purpose-built connectivity solutions. By ... We welcome both recent graduates with strong, directly relevant project, research, or internship ...

next page

Showing results 1-20

Internship Scale Ai information

What cities are hiring for Internship Scale Ai jobs? Cities with the most Internship Scale Ai job openings:
What states have the most Internship Scale Ai jobs? States with the most job openings for Internship Scale Ai jobs include:

AI Systems Performance Engineer - New Graduate

SambaNova Systems

San Jose, CA

Other

Posted 18 days ago


Job description

About The Role

We are seeking a talented and highly motivated AI Systems Performance Engineer to bring up and optimize state-of-the-art foundation models on SambaNova's reconfigurable dataflow platform.

You'll work hands-on with advanced AI models - such as DeepSeek, GLM, Kimi, GPT OSS, Llama, Qwen, and other frontier architectures - and learn how modern AI systems achieve high throughput, low latency, and efficient large-scale inference.

In this role, you'll work at the intersection of machine learning and computer systems, collaborating with engineers across model, compiler, runtime, and hardware teams. This is an ideal opportunity for a new graduate who is passionate about understanding how AI models execute on real hardware and wants to help build the next generation of high-performance AI systems.

Responsibilities
  • Bring up cutting-edge foundation models, including LLMs and multimodal models, on the SambaNova platform through the SambaNova software stack.
  • Analyze and profile model execution to identify performance bottlenecks across model, compiler, runtime, and hardware layers.
  • Optimize AI workloads for throughput, latency, memory efficiency, and scalability.
  • Collaborate with machine learning, compiler, runtime, and hardware engineers to develop high-performance AI applications.
  • Explore and integrate new techniques in model architecture, quantization, scheduling, caching, and memory optimization.
  • Develop tools, benchmarks, and performance analysis methodologies for large-scale AI inference.
  • Investigate new model architectures and translate research advances into efficient implementations on production AI systems.
  • Contribute ideas for dataflow, scheduling, and system optimizations for both single-node and distributed inference.
Basic Qualifications
  • Bachelor's or Master's degree in computer science, electrical engineering, computer engineering, or a related technical field (e.g., applied mathematics, physics, or statistics), completed or expected before the start date.
  • Strong programming skills in Python, C++, or a similar programming language.
  • Solid foundations in algorithms, data structures, computer architecture, operating systems, or parallel computing.
  • Familiarity with deep learning and at least one major ML framework, such as PyTorch, TensorFlow, or JAX.
  • Strong analytical and problem-solving skills, with an interest in understanding and optimizing system performance.
  • Ability and enthusiasm to learn across machine learning, software systems, and hardware.
Preferred Qualifications
  • Coursework, research, internship, or project experience in machine learning systems, computer architecture, compilers, distributed systems, or high-performance computing.
  • Hands-on experience with LLMs, multimodal models, or transformer architectures.
  • Familiarity with model inference, KV cache, batching, quantization, or distributed execution.
  • Experience with GPU or accelerator programming using CUDA, Triton, OpenCL, or similar technologies.
  • Familiarity with frameworks such as vLLM, DeepSpeed, Megatron, or TensorRT.
  • Understanding of memory hierarchy, caching, parallelism, or scheduling.
  • Experience profiling and optimizing the performance of software or ML workloads.
  • Research publications, open-source contributions, programming competitions, or technically challenging personal projects are a plus.

We value strong technical fundamentals, curiosity, and the ability to learn quickly. Prior production experience with large-scale AI systems is not required.