← nvidia / Senior Technical Product Manager - GPU Direct Storage
cover_letter / art_hZ-m96y1f4w
role
model
anthropic/claude-sonnet-4.6
created
2026-05-20T22:38
Cover letter
Dear NVIDIA GPUDirect Storage Hiring Team,
NVIDIA's CUDA platform sits at the foundation of nearly every meaningful advance in AI, HPC, and scientific computing today — and GPUDirect Storage represents one of the most consequential bottlenecks left to solve: getting data from storage to GPU memory without the CPU becoming the limiting factor. That problem is deeply familiar to me. When I scaled Intuit's ICE platform to 675M+ engagements in FY23 and drove a rSocket migration that pushed throughput from 6K to 50K TPS while supporting ~1.5M concurrent connections at sub-25ms TP99, the recurring lesson was that infrastructure bottlenecks — not compute — are what stall real-world AI and data-intensive workloads. GPUDirect Storage attacks exactly that class of problem at the GPU level, and that is why this role stands out.
**Technical Foundation**
My technical credibility spans from low-level systems work to applied ML infrastructure. In 2004, I hand-coded a neural network in C++ with custom backpropagation through time (BPTT) for protein structure prediction — work that was accepted at NeurIPS 2014 after further development. In 2026, I rewrote that system as a production ML platform in PyTorch spanning 413 parameters to 8B (a 19-million-fold scale increase), with MLflow experiment tracking, Optuna hyperparameter optimization, FastAPI serving, and Docker orchestration across 6 containers. That arc — from hand-writing gradient descent in C++ to orchestrating multi-container ML pipelines — reflects the kind of depth-first technical engagement I bring to product work.
More directly relevant to this role: I built an RL post-training workbench that benchmarks GRPO, DPO, PPO, DAPO, and 8 additional algorithms across TRL, VeRL, OpenRLHF, and NeMo RL, with GPU passthrough in Docker containers and live SSE metric streaming on Apple Silicon (MPS) and CUDA. Designing that system required understanding GPU memory constraints, throughput characteristics, and framework-level differences in how training loops interact with hardware — exactly the kind of reasoning that applies to GPUDirect Storage's value proposition in AI training pipelines. I also built aeval, a local-first model evaluation platform with a FastAPI orchestrator, TimescaleDB, Redis job queue, and Ollama integration — demonstrating sustained ability to architect and ship production-grade developer infrastructure end to end.
**Why This Role**
My career has consistently sat at the intersection of developer-facing platform infrastructure and technically demanding customers — from owning Splunk's Search Service (Go microservices, SPL/SPL2) and Search Catalog (PostgreSQL metadata service) to leading Intuit's developer framework strategy across 9 programming languages and 30+ product SKUs. GPUDirect Storage and cuFile occupy a similar position: a deep systems technology that must be made accessible, well-documented, and strategically positioned for a developer audience that ranges from HPC researchers to AI infrastructure engineers. That is the product management motion I know best.
What excites me specifically about this role is the combination of roadmap ownership, customer engagement, and thought leadership. The JD calls for creating technical content — blogs, white papers, webinars, tutorials — and I have done exactly that: I was a DeveloperWeek 2022 speaker, a Splunk .conf18 and .conf19 speaker, and I am a forthcoming O'Reilly author on Data Analytics and Cloud Project Management. Translating GPUDirect Storage's performance characteristics into compelling, technically precise content for AI and HPC practitioners is a task I am well-positioned to execute. I am also drawn to the customer engagement dimension: at Splunk, I led a query performance optimization initiative with a beta enterprise customer that achieved up to 10x performance improvements in Splunk Cloud Services search — the kind of tight, trust-based technical collaboration with demanding customers that NVIDIA's developer ecosystem requires.
**Selected Relevant Experience**
- **Scaled Intuit ICE platform to 675M+ engagements (FY23)** across QuickBooks, TurboTax, Mint, Mailchimp, and Credit Karma; drove rSocket migration increasing throughput from 6K to 50K TPS, supporting ~1.5M concurrent connections at sub-25ms TP99 — directly analogous to the high-throughput, low-latency infrastructure concerns central to GPUDirect Storage.
- **Built RL post-training workbench** with GPU Docker passthrough, CUDA/MPS support, and live metric streaming across 12 RL algorithms and 4 frameworks (TRL, VeRL, OpenRLHF, NeMo RL) — hands-on experience with GPU training infrastructure and the data pipeline bottlenecks GPUDirect Storage addresses.
- **Delivered Intuit ICE Self-Service platform** (DevPortal, GitOps config, ICE Playground), reducing developer onboarding from 2–3 weeks to minutes in pre-prod — demonstrated ability to ship developer-facing platform products that drive measurable adoption.
- **Extended Java and Python SDK Starter Kits** with scaffolding templates, build configurations (Gradle/Maven), testing frameworks, and CI/CD integration — experience owning SDK and developer tooling products at scale.
- **Led Splunk Search Service and SPL/SPL2 product ownership** (Go microservices, PostgreSQL metadata service), delivering Scheduler Service end-to-end in ~4 months and achieving up to 10x query performance improvements with an enterprise beta customer.
- **Conducted enterprise-wide Service Language Assessment** across 9 languages at Intuit, analyzing usage data and developer feedback to inform strategic investment decisions presented to the CTO — the kind of cross-functional, data-driven strategic work the GPUDirect Storage roadmap role demands.
- **NeurIPS 2014 published researcher** (protein structure prediction via artificial neural networks); original 2004 system hand-coded in C++ with custom BPTT — establishing foundational technical credibility in ML systems and scientific computing.
**Closing**
NVIDIA's mission to accelerate computing is not abstract to me — it is the substrate on which every AI workload I have built or benchmarked depends. GPUDirect Storage removes one of the last major CPU-bound bottlenecks in GPU-accelerated pipelines, and the opportunity to own that product's roadmap, customer engagement, and technical narrative at NVIDIA is one I take seriously. I would welcome the opportunity to discuss how my background in developer platform infrastructure, applied ML systems, and technical product leadership maps to what your team is building.
Thank you for your consideration.
---
**O. Felix Amoruwa**
famoruwa@berkeley.edu | 909-731-9011 | felixamoruwa.info