← nvidia / Principal Product Manager
tailored_resume_v2 / art_3WYo9CdTDjw
↓ Download .docx ↓ Download .pdf PDF requires LibreOffice installed
What changed for nvidia
| change | why it matters |
|---|---|
| Summary rewritten to lead with '15+ years building infrastructure platforms and orchestration systems at scale' and embed SLO ownership, failure attribution, operator UX, and GPU infrastructure credentials | JD's first hard requirement is 15+ years PM in infrastructure/platform/MLOps; summary must immediately establish fit at the Principal level |
| Intuit role retitled to 'Platform Infrastructure & Workflow Orchestration' and reordered to lead experience section | Intuit is the strongest proof point for distributed systems at scale, SLO ownership, and operator-facing automation — directly maps to DGX Cloud break-fix PM scope |
| Intuit bullet 1 leads with 675M engagements / 50K TPS / sub-25ms TP99 fleet availability framing | JD defines success as time-to-healthy and fleet availability; enterprise-scale metrics must appear in first bullet |
| Intuit Drift Detection bullet reframed as 'audit trails and automated repair workflows' using JD language | Drift Detection is the closest analog to break-fix automation with audit trails — JD explicitly requires both |
| Splunk role retitled to 'Search Orchestration & Distributed Systems' and 10x performance bullet leads | JD requires distributed systems expertise and reliability engineering; 10x improvement demonstrates blast radius management instincts |
| Kaiser role retitled to 'Infrastructure Reliability' and Redis/fault-tolerance bullet elevated | JD preferred qual includes reliability engineering background; Redis caching for fault tolerance is the strongest proof point here |
| StreamIO role condensed to 2 bullets focused on OpenClaw multi-agent orchestration and MCP SDK | JD preferred qual explicitly calls out 'agentic AI workflow software'; OpenClaw is the strongest match |
| RL Workbench project moved to lead the projects section | GPU Docker passthrough and multi-framework benchmarking is the strongest proof of GPU infrastructure experience — a JD preferred qualification |
| aeval project reframed around 'automated safety gates' and 'automation confidence thresholds' | JD requires defining automation confidence thresholds and human-in-the-loop intervention points; aeval's safety gates are the closest technical analog |
| AutoEval project reframed as 'detection-to-resolution pipeline with audit trails' | JD explicitly requires driving integration from failure detection to resolution — AutoEval's 72hr→4min cycle reduction demonstrates this end-to-end |
| IBM bullet reframed to emphasize 'failure attribution workflows' and 'blast radius management' | JD requires track record owning products with real-world operational consequences; IBM escalation resolution maps to this requirement |
| Fintellect condensed to 1 bullet focused on multi-provider LLM orchestration with fallback routing as reliability-first agentic workflow | Lower relevance role; 1 bullet preserves presence while freeing space for higher-relevance content |
JD analysis (20 key phrases)
Key phrases: break-fix automationAI factoryresilient automationworkflow orchestrationoperator UXrepair queuesaudit trailshuman-in-the-loopautomation confidence thresholdstime-to-draintime-to-healthyfleet availabilitySLOblast radiusfailure attributionRMA processesNCP operatorsSRE teamsagentic AI workflowself-heal
Hard requirements:
- 15+ years PM experience in infrastructure, platform, or MLOps
- BS or MS in CS, Engineering, or related technical field
- Distributed systems and workflow orchestration expertise
- Safety tradeoffs in automation (blast radius awareness)
- Operator UX for complex system state under pressure
- Alignment across engineering, SRE, and external vendor partners
- SLO definition and metrics framework ownership
Preferred qualifications:
- GPU infrastructure or datacenter operations experience
- RMA logistics and vendor SLA oversight at scale
- Reliability engineering, SLO build, chaos/fault-injection testing
- Cloud service provider or hyperscaler infrastructure background
- Agentic AI workflow software experience
Per-role mapping (9 roles scored)
| role | score | reframe angle | JD phrases that map |
|---|---|---|---|
| Intuit — Staff PM Developer Frameworks & Platform Infrastructure | 5/5 | Large-scale platform infrastructure PM with SLO ownership, drift detection/remediation, and distributed systems at hyperscaler throughput | workflow orchestration, fleet availability, SLO, blast radius, audit trails, operator UX, distributed systems |
| Splunk — Senior PM Search Orchestration | 4/5 | Distributed search orchestration PM with reliability focus, SLA management, and enterprise-scale performance optimization | workflow orchestration, distributed systems, SLO, failure attribution, operator UX |
| Kaiser Permanente — SOA Technical PM | 3/5 | Enterprise infrastructure reliability PM with datacenter capacity planning and fault-tolerance architecture | fleet availability, SLO, blast radius, distributed systems |
| StreamIO AI — Founder & CEO | 3/5 | Agentic AI workflow orchestration builder with multi-agent coordination and production deployment | agentic AI workflow, workflow orchestration, human-in-the-loop |
| Fintellect AI — Founder & CEO | 2/5 | AI orchestration with reliability routing | workflow orchestration, agentic AI workflow |
| IBM — Software Engineer BI Products | 2/5 | Enterprise escalation and root-cause resolution — early foundation for operational reliability | failure attribution, blast radius |
| RL Workbench Project | 4/5 | GPU infrastructure and MLOps tooling — directly relevant to AI factory context | AI factory, GPU infrastructure, MLOps, fleet availability |
| aeval Project | 3/5 | Automated evaluation with safety gates — analogous to automation confidence thresholds and human-in-the-loop criteria | automation confidence thresholds, human-in-the-loop, audit trails |
| AutoEval Project | 3/5 | Automated repair-cycle analogy — detection to resolution pipeline with confidence scoring | failure attribution, automation confidence thresholds, time-to-healthy |
Tailored summary
Technical Product Leader with 15+ years building infrastructure platforms and orchestration systems at scale — scaling distributed services to 675M+ engagements and 50K TPS at Intuit, owning search orchestration microservices at Splunk, and architecting agentic AI workflow frameworks in production today. Proven track record defining SLOs, driving failure attribution to automated resolution, and translating complex system state into operator UX that engineers can act on under pressure. NeurIPS published researcher; hands-on GPU infrastructure experience benchmarking RL frameworks across CUDA/MPS with Docker passthrough.