San Francisco, CA · On-site · Full-time Compensation: $150,000–$250,000 + 0%–1% equity
About the Company
Our client does LLM interpretability and context-optimization research, building custom machine-learning models that analyze and compress token contexts before they reach the underlying model. The result is roughly a 50% inference-cost reduction, lower latency, and measurably higher accuracy for the enterprises and scale-ups integrating LLMs into their products. About seven months old, it already serves roughly 1,000 customers and is well-backed by top-tier investors and notable operators.
Founded 2025 · 1–10 people (Seed) · Industry: AI Tools
The Role
Own the full multi-region GPU infrastructure stack end to end as the sole infra hire — global low-latency serving, multi-cloud and on-prem deployments, reliability, and cost efficiency — for a seed-stage LLM context-compression company. In-person in SF at a high-intensity pace.
Tech stack: AWS, GCP, Base10, Terraform, Docker, CI/CD, GPU/ML inference infrastructure, AWS Marketplace.
What you'll be doing
- Own the cloud systems serving the compression API end to end
- Build and operate global, low-latency, high-throughput GPU ML inference infrastructure
- Work across AWS, Terraform, Docker, and CI/CD
- Continuously improve and research infrastructure solutions
Requirements
- Has built and operated production infrastructure at a startup or larger company
- Learns new solutions and technologies quickly
- Based in or willing to relocate to San Francisco to work in person at the hacker house
- Willing to work startup hours in a high-intensity (9am–9pm, six-day) environment
Nice to Haves
- A quick learner who grasps products and systems fast
- Experience building for performance and reliability at scale
- A research-and-product-focused mindset
- A high-ownership mentality
- A startup-minded operator who prioritizes learning and growth over work-life balance
- GPU-infrastructure experience in production, or a background at an infrastructure company
- First infra-hire experience at a startup
Why Join
- Sole infra owner, full stack, day one: own every layer of a multi-region GPU stack (AWS, GCP, Base10, on-prem), sitting directly in the critical path of ~1,000 customers
- Strong backing and early traction: ~1,000 customers within seven months, well-funded by top-tier investors
- Comp and lifestyle support for high output: $150K–$250K + up to 1% equity, with housing and food provided at the SF hacker house, visa sponsorship, laundry/cleaning, meal delivery, and health/dental
Details
- Location — San Francisco, CA
- Work policy — On-site, high intensity (six days/week)
- Compensation — $150,000–$250,000 + 0%–1% equity
- Visa sponsorship — Available (H-1B, O-1, OPT)
- Employment type — Full-time