← nvidia / Senior Technical Product Manager - GPU Direct Storage
brief / art_CRoMHqtMEdA
role
model
anthropic/claude-sonnet-4.6
created
2026-05-20T22:39
Company snapshot
NVIDIA is the dominant GPU platform company, best known for inventing the CUDA parallel computing ecosystem and the GeForce/Tesla/H100 GPU lines that now underpin virtually all AI training and inference infrastructure. Over the last 12–24 months NVIDIA has expanded aggressively into full-stack AI infrastructure: the Hopper (H100) and Blackwell (B100/B200) GPU generations, NIM microservices, and the CUDA-X library ecosystem. GPUDirect Storage (GDS) and the cuFile API are production NVIDIA technologies that enable direct DMA paths between NVMe/NFS storage and GPU memory, eliminating CPU-bounce-buffer overhead — critical for AI dataset ingestion, HPC checkpointing, and large-scale inference. NVIDIA's engineering culture is deeply technical; PMs are expected to write PRDs, engage directly with HPC/AI customers, and produce credible technical content. Specific recent internal project names and personnel are not publicly confirmed and are not stated here.
Team stack
Based on the JD and public NVIDIA documentation: core product is GPUDirect Storage (GDS) + cuFile C/C++ library running on Linux (RHEL, Ubuntu); CUDA runtime and NVIDIA driver stack (likely CUDA 12.x); storage backends include NVMe-oF, local NVMe, Lustre, GPFS/IBM Spectrum Scale, WekaFS, and NFS over RDMA (likely, based on known GDS partner ecosystem); RDMA/InfiniBand networking (likely, given HPC positioning); container/orchestration layer likely includes Docker + Kubernetes + NVIDIA GPU Operator; benchmark and profiling tooling likely includes Nsight Systems, nvtop, and custom throughput harnesses; CI/CD likely Jenkins or GitLab CI (internal, unconfirmed); customer environments span on-prem HPC clusters and cloud (AWS, GCP, Azure) — cloud angle confirmed by JD mention of 'cloud-based or HPC environments'.
Likely questions (10)
| area | question | why |
|---|---|---|
| domain | Walk us through how GPUDirect Storage eliminates the CPU bounce buffer. What are the performance implications for a large-scale AI training job loading data from NVMe? | The JD explicitly lists 'strong technical expertise in GPUDirect Storage, cuFile, or closely related technologies' as a hard requirement. This is the core domain knowledge gate. |
| system_design | Design a data-ingestion pipeline for a 10,000-GPU training cluster that must sustain >1 TB/s aggregate read throughput from a distributed file system using GDS. What are the bottlenecks and how do you instrument them? | JD calls out 'scientific computing, AI workloads, and HPC environments' and 'large-scale data management.' This tests whether the candidate can reason about GDS at cluster scale. |
| domain | How does the cuFile API differ from POSIX I/O from the developer's perspective? What are the integration pain points you'd expect developers to hit, and how would you address them in the SDK? | JD emphasizes 'customer engagement' and 'developer pain points.' cuFile is the programmatic surface of GDS; understanding its ergonomics is essential for roadmap and DevRel work. |
| behavioral | Tell me about a time you drove a developer platform from low adoption to significant scale. What metrics did you use, and what were the key levers? | Candidate's Intuit ICE platform scaled to 675M engagements and 50K TPS — directly maps to JD's 'customer engagement' and 'go-to-market strategy' requirements. |
| behavioral | Describe a situation where you had to synthesize conflicting requirements from HPC researchers, cloud platform teams, and storage hardware vendors into a single coherent roadmap. | JD requires collaboration with 'architects, engineers, researchers' and cross-functional GTM teams. GDS sits at the intersection of storage vendors, cloud providers, and AI/HPC users. |
| coding | You're reviewing a developer's cuFile integration and their benchmark shows only 40% of theoretical NVMe bandwidth is being utilized. Walk me through your debugging methodology — what would you check first? | JD lists 'direct experience with CUDA, GPU programming, and large-scale data management' as a differentiator. This tests practical performance debugging depth. |
| system_design | How would you design a benchmarking framework to compare GDS performance across different storage backends (local NVMe, NVMe-oF, Lustre, WekaFS) in a reproducible way for use in white papers and customer collateral? | JD explicitly asks for 'thought leadership: blogs, white papers, webinars' and the candidate has directly built a multi-framework benchmarking workbench (RL Workbench with GPU Docker passthrough). |
| domain | What is your understanding of RDMA and how does it relate to GPUDirect Storage? How would you explain the difference between GPUDirect RDMA and GPUDirect Storage to a developer new to the ecosystem? | GDS relies on RDMA for NVMe-oF paths; the JD's HPC focus means candidates must understand the full DMA/RDMA stack and be able to communicate it clearly. |
| culture | NVIDIA PMs are expected to produce technical content — blogs, white papers, conference talks — not just PRDs. Tell me about technical content you've created and the impact it had on developer adoption or customer understanding. | JD explicitly lists 'thought leadership' as a core responsibility. Candidate has DeveloperWeek and Splunk .conf speaking history plus published blog evidence. |
| behavioral | Give an example of a time you used quantitative data (telemetry, benchmarks, usage analytics) to make a prioritization decision on a platform product. What was the data, what did you decide, and what was the outcome? | JD requires 'refining product strategy' and 'customer engagement.' Candidate's Intuit experience with BigQuery/SQL telemetry across 20 mobile apps is directly relevant. |
Talking points
- Platform scale with hard numbers: At Intuit I owned the ICE developer platform end-to-end — grew engagements 275% YoY to 675M in FY23, scaled throughput from 6K to 50K TPS via rSocket migration supporting ~1.5M concurrent connections at sub-25ms TP99. I know what it takes to take a developer infrastructure product from early adoption to enterprise scale, and I know how to instrument it to prove the value.
- Hands-on GPU and ML infrastructure depth: I built an RL post-training workbench that benchmarks GRPO/DPO across TRL, VeRL, OpenRLHF, and NeMo RL with GPU Docker passthrough and live SSE metric streaming on Apple Silicon MPS and CUDA — 12 algorithms, standardized throughput/memory/convergence benchmarking. I understand GPU memory hierarchies, compute/memory bandwidth tradeoffs, and the data-pipeline bottlenecks that GDS is designed to solve.
- Developer SDK and tooling ownership: At Intuit I extended Java and Python SDK Starter Kits with scaffolding, Gradle/Maven build configs, testing frameworks, and CI/CD integration; I also built the ICE Self-Service DevPortal and reduced developer onboarding from 2–3 weeks to under 24 hours. Translating that to cuFile means I can own the full developer experience — API ergonomics, documentation, playground environments, and adoption metrics.
- Technical content and public speaking: I've spoken at DeveloperWeek 2022 and Splunk .conf18/.conf19, and I've published blogs on multi-framework ML benchmarking, multi-agent orchestration, and AI evaluation platforms. I can produce the white papers, webinars, and technical collateral the JD calls out — not as a side task but as a core product motion.
- NeurIPS-published researcher with HPC roots: My 2014 NeurIPS paper on neural networks for protein secondary structure prediction, combined with computational biology work at Lawrence Berkeley National Lab using Monte Carlo and AI algorithms on RNA/DNA sequences, gives me credibility in the scientific computing and HPC communities that are GDS's primary customer base — I can engage researchers and HPC architects as a peer, not just a PM.