Job Summary:
Positron AI specializes in developing custom hardware systems to accelerate AI inference, and they are seeking a Senior Software Engineer to contribute to the development of high-performance software for executing open-source large language models on their custom appliance. The role involves designing and implementing high-performance inference software, optimizing C++-based libraries, and collaborating with engineers to ensure efficient data movement between CPUs and FPGAs.
Responsibilities:
• Design and implement high-performance inference software for LLMs on custom hardware.
• Develop and optimize C++-based libraries that efficiently utilize SIMD instructions, threading, and memory hierarchy.
• Work closely with FPGA and systems engineers to ensure efficient data movement and computational offloading between x86 CPUs and FPGAs.
• Optimize model execution via low-level optimizations, including vectorization, cache efficiency, and hardware-aware scheduling.
• Contribute to performance profiling tools and methodologies to analyze execution bottlenecks at the instruction and data flow levels.
• Apply NUMA-aware memory management techniques to optimize memory access patterns for large-scale inference workloads.
• Implement ML system-level optimizations such as token streaming, KV cache optimizations, and efficient batching for transformer execution.
• Collaborate with ML researchers and software engineers to integrate model quantization techniques, sparsity optimizations, and mixed-precision execution.
• Ensure all code contributions include unit, performance, acceptance, and regression tests as part of a continuous integration-based development process.
Qualifications:
Required:
• 7+ years of professional experience in C++ software development, with a focus on performance-critical applications.
• Strong understanding of C++ templates and modern memory management.
• Hands-on experience with SIMD programming (AVX-512, SSE, or equivalent) and intrinsics-based vectorization.
• Experience in high-performance computing (HPC), numerical computing, or ML inference optimization.
• Experience with ML model execution optimizations, including efficient tensor computations and memory access patterns.
• Knowledge of multi-threading, NUMA architectures, and low-level CPU optimization.
• Proficiency with systems-level software development, profiling tools (perfetto, VTune, Valgrind), and benchmarking.
• Experience working with hardware accelerators (FPGAs, GPUs, or custom ASICs) and designing efficient software-hardware interfaces.
Preferred:
• Familiarity with LLVM/Clang or GCC compiler optimizations.
• Experience in LLM quantization, sparsity optimizations, and mixed-precision computation.
• Knowledge of distributed inference techniques and networking optimizations.
• Understanding of graph partitioning and execution scheduling for large-scale ML models.
Company:
Positron delivers vendor freedom and faster inference for both enterprises and research teams, by allowing them to use hardware and software explicitly designed from the ground up for generative and large language models (LLMs). Founded in 2023, the company is headquartered in Reno, USA, with a team of 11-50 employees. The company is currently Early Stage.