jobsearch v0.0.1

← nvidia / Senior Technical Product Manager - GPU Direct Storage

brief / art_CRoMHqtMEdA

role
nvidia / Senior Technical Product Manager - GPU Direct Storage
model
anthropic/claude-sonnet-4.6
created
2026-05-20T22:39

Company snapshot

NVIDIA is the dominant GPU platform company, best known for inventing the CUDA parallel computing ecosystem and the GeForce/Tesla/H100 GPU lines that now underpin virtually all AI training and inference infrastructure. Over the last 12–24 months NVIDIA has expanded aggressively into full-stack AI infrastructure: the Hopper (H100) and Blackwell (B100/B200) GPU generations, NIM microservices, and the CUDA-X library ecosystem. GPUDirect Storage (GDS) and the cuFile API are production NVIDIA technologies that enable direct DMA paths between NVMe/NFS storage and GPU memory, eliminating CPU-bounce-buffer overhead — critical for AI dataset ingestion, HPC checkpointing, and large-scale inference. NVIDIA's engineering culture is deeply technical; PMs are expected to write PRDs, engage directly with HPC/AI customers, and produce credible technical content. Specific recent internal project names and personnel are not publicly confirmed and are not stated here.

Team stack

Based on the JD and public NVIDIA documentation: core product is GPUDirect Storage (GDS) + cuFile C/C++ library running on Linux (RHEL, Ubuntu); CUDA runtime and NVIDIA driver stack (likely CUDA 12.x); storage backends include NVMe-oF, local NVMe, Lustre, GPFS/IBM Spectrum Scale, WekaFS, and NFS over RDMA (likely, based on known GDS partner ecosystem); RDMA/InfiniBand networking (likely, given HPC positioning); container/orchestration layer likely includes Docker + Kubernetes + NVIDIA GPU Operator; benchmark and profiling tooling likely includes Nsight Systems, nvtop, and custom throughput harnesses; CI/CD likely Jenkins or GitLab CI (internal, unconfirmed); customer environments span on-prem HPC clusters and cloud (AWS, GCP, Azure) — cloud angle confirmed by JD mention of 'cloud-based or HPC environments'.

Likely questions (10)

areaquestionwhy
domain Walk us through how GPUDirect Storage eliminates the CPU bounce buffer. What are the performance implications for a large-scale AI training job loading data from NVMe? The JD explicitly lists 'strong technical expertise in GPUDirect Storage, cuFile, or closely related technologies' as a hard requirement. This is the core domain knowledge gate.
system_design Design a data-ingestion pipeline for a 10,000-GPU training cluster that must sustain >1 TB/s aggregate read throughput from a distributed file system using GDS. What are the bottlenecks and how do you instrument them? JD calls out 'scientific computing, AI workloads, and HPC environments' and 'large-scale data management.' This tests whether the candidate can reason about GDS at cluster scale.
domain How does the cuFile API differ from POSIX I/O from the developer's perspective? What are the integration pain points you'd expect developers to hit, and how would you address them in the SDK? JD emphasizes 'customer engagement' and 'developer pain points.' cuFile is the programmatic surface of GDS; understanding its ergonomics is essential for roadmap and DevRel work.
behavioral Tell me about a time you drove a developer platform from low adoption to significant scale. What metrics did you use, and what were the key levers? Candidate's Intuit ICE platform scaled to 675M engagements and 50K TPS — directly maps to JD's 'customer engagement' and 'go-to-market strategy' requirements.
behavioral Describe a situation where you had to synthesize conflicting requirements from HPC researchers, cloud platform teams, and storage hardware vendors into a single coherent roadmap. JD requires collaboration with 'architects, engineers, researchers' and cross-functional GTM teams. GDS sits at the intersection of storage vendors, cloud providers, and AI/HPC users.
coding You're reviewing a developer's cuFile integration and their benchmark shows only 40% of theoretical NVMe bandwidth is being utilized. Walk me through your debugging methodology — what would you check first? JD lists 'direct experience with CUDA, GPU programming, and large-scale data management' as a differentiator. This tests practical performance debugging depth.
system_design How would you design a benchmarking framework to compare GDS performance across different storage backends (local NVMe, NVMe-oF, Lustre, WekaFS) in a reproducible way for use in white papers and customer collateral? JD explicitly asks for 'thought leadership: blogs, white papers, webinars' and the candidate has directly built a multi-framework benchmarking workbench (RL Workbench with GPU Docker passthrough).
domain What is your understanding of RDMA and how does it relate to GPUDirect Storage? How would you explain the difference between GPUDirect RDMA and GPUDirect Storage to a developer new to the ecosystem? GDS relies on RDMA for NVMe-oF paths; the JD's HPC focus means candidates must understand the full DMA/RDMA stack and be able to communicate it clearly.
culture NVIDIA PMs are expected to produce technical content — blogs, white papers, conference talks — not just PRDs. Tell me about technical content you've created and the impact it had on developer adoption or customer understanding. JD explicitly lists 'thought leadership' as a core responsibility. Candidate has DeveloperWeek and Splunk .conf speaking history plus published blog evidence.
behavioral Give an example of a time you used quantitative data (telemetry, benchmarks, usage analytics) to make a prioritization decision on a platform product. What was the data, what did you decide, and what was the outcome? JD requires 'refining product strategy' and 'customer engagement.' Candidate's Intuit experience with BigQuery/SQL telemetry across 20 mobile apps is directly relevant.

Talking points