1

Self Decode Jobs (NOW HIRING)

GenBio AI develops multiscale foundation models to decode and simulate human biology. Our team is ... Motivated and self-driven with the ability to operate with partial and incomplete descriptions of ...

Be Seen First

... self-starter who thrives on high autonomy, rapid delivery, and fast-paced environments. You don't just wait for instructions--you decode the vision, map out the systems, and execute flawlessly. Key ...

A self-starter, who'll take ideas from concept to execution, navigating feedback like a pro What ... Present creative ideas directly to clients with confidence; decode ambiguous client feedback into ...

A self-starter, who'll take ideas from concept to execution, navigating feedback like a pro What ... Present creative ideas directly to clients with confidence; decode ambiguous client feedback into ...

GenBio AI develops multiscale foundation models to decode and simulate human biology. Our team is ... Motivated and self-driven with the ability to operate with partial and incomplete descriptions of ...

GenBio AI develops multiscale foundation models to decode and simulate human biology. Our team is ... Motivated and self-driven with the ability to operate with partial and incomplete descriptions of ...

GenBio AI develops multiscale foundation models to decode and simulate human biology. Our team is ... Motivated and self-driven with the ability to operate with partial and incomplete descriptions of ...

A self-starter, who'll take ideas from concept to execution, navigating feedback like a pro What ... Present creative ideas directly to clients with confidence; decode ambiguous client feedback into ...

A self-starter, who'll take ideas from concept to execution, navigating feedback like a pro What ... Present creative ideas directly to clients with confidence; decode ambiguous client feedback into ...

Senior Automotive Hardware Architect

Detroit, MI · On-site

$97K - $124K/yr

... decode/encode, Image processing, Audio processing, Camera & Display, Networking, and advanced ... Self-starter with the ability to work independently given minimal supervision * Ability to work in ...

Showing results 21-40

Self Decode information

See salary details

$46K

$115.4K

$172.5K

How much do self decode jobs pay per year?

As of Aug 21, 2026, the average yearly pay for self decode in the United States is $115,438.00, according to ZipRecruiter salary data. Most workers in this role earn between $80,500.00 and $134,000.00 per year, depending on experience, location, and employer.
Infographic showing various Self Decode job openings in the United States as of August 2026, with employment types broken down into 1% As Needed, 78% Full Time, 18% Part Time, and 3% Contract. Highlights an 91% Physical, 2% Hybrid, and 7% Remote job distribution, with an average salary of $115,438 per year, or $55.5 per hour.

Member of Technical Staff (AI Inference Engineer)

Perplexity

San Francisco, CA • On-site

Full-time

Re-posted 20 days ago


Job description

Job Summary:
Perplexity is a company that builds and runs the inference engine behind every query. They are seeking a Member of Technical Staff (AI Inference Engineer) to support transformer-based models, develop a Rust-native serving runtime, and optimize performance and reliability within their inference infrastructure.
Responsibilities:
• New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway.
• GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow.
• Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic.
• Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernel interleaving.
• Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents.
Qualifications:
Required:
• 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
• Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar). Any other deep systems programming experience is a plus.
• You understand modern LLM architectures and are able to bring them up reliably in a production environment.
• You've built and operated production distributed systems under real load - ideally performance-critical ones.
• Comfortable working across languages and layers: Rust for the serving runtime, Python for model code, CUDA/CuteDSL for kernels.
• You own problems end-to-end. You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on Friday.
• Self-directed. You do well in fast-moving environments where the path forward isn't laid out for you.
• Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).
• Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores).
• Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation).
Preferred:
• Good if you touched any of ML compilers and framework internals: PyTorch internals, torch.compile, custom operators.
• Good if you touched any of Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism.
• Good if you touched any of Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving.
• Good if you touched any of Profiling and debugging tools: Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis.
• Good if you touched any of Container orchestration: Kubernetes, GPU scheduling, autoscaling inference workloads.
Company:
Perplexity is an AI-powered platform that retrieves, analyzes information from the web to deliver structured answers with cited sources. Founded in 2022, the company is headquartered in San Francisco, USA, with a team of 201-500 employees. The company is currently Growth Stage.