← nvidia / Principal Product Manager
cover_letter / art_qNrvZwKLnWQ
Cover letter
Dear NVIDIA Hiring Team,
NVIDIA's AI factory vision — converting tokens to intelligence at scale — sits at the center of how the next decade of compute will be built and operated. The break-fix automation challenge you're solving is not a peripheral concern; it is the operational spine that determines whether AI infrastructure delivers on its SLA promises or collapses under its own complexity. My interest in this role is grounded in a specific arc: I have spent the last several years building systems where automation must be trusted enough to act, but transparent enough that humans can intervene confidently — from platform infrastructure at Intuit serving 675M+ engagements to agentic orchestration frameworks I designed from scratch.
**Technical and AI/ML Foundation**
My technical foundation spans both the systems layer and the AI layer that increasingly governs it. At Intuit, I owned the ICE platform — a developer framework and infrastructure layer serving QuickBooks, TurboTax, Mint, Mailchimp, and Credit Karma simultaneously. Scaling that system from 6K to 50K TPS via rSocket migration, supporting approximately 1.5M concurrent connections at sub-25ms TP99, required exactly the kind of blast-radius thinking the JD describes: every configuration change, every migration, every deprecation had real operational consequences for downstream product teams. I built the MSaaS Drift Detection program — a Java JAR library that scanned Git repositories for configuration drift — because I understood that undetected divergence in distributed systems is a silent reliability risk. That instinct maps directly to the failure attribution and automated repair problem NVIDIA is solving.
On the agentic AI side, I built OpenClaw, a multi-agent orchestration framework with a gateway protocol, subagent delegation, profile management, and session switching — enabling coordinated AI agent workflows across multiple industry verticals. I also built aeval, a local-first model evaluation platform with FastAPI orchestration, TimescaleDB for time-series metric storage, a Redis job queue, and automated safety gates with regression detection — a system designed around the principle that automation must earn confidence through measurable, auditable evidence before it acts. My RL Workbench benchmarks 12 algorithms (PPO, GRPO, DAPO, DPO, and others) across TRL, VeRL, OpenRLHF, and NeMo RL with GPU Docker passthrough — work that required me to reason carefully about GPU resource allocation, container isolation, and throughput/memory tradeoffs at the framework level.
My research foundation includes a NeurIPS 2014 accepted paper on artificial neural networks for protein secondary structure prediction, and my original 2004 system was a hand-coded neural network in C++ with custom backpropagation through time — context that grounds my AI work in first-principles understanding rather than API consumption.
**Why This Role**
The break-fix automation problem at NVIDIA AI Factory is one of the most consequential reliability engineering challenges in the industry right now — and it is precisely the intersection of distributed systems, operator UX, and agentic automation where my background converges.
What specifically draws me to this role is the human-in-the-loop design challenge: defining automation confidence thresholds and blocking criteria that let the system act fast without creating new failure modes. I have navigated this tradeoff in production — at Intuit, the ICE Self-Service platform reduced developer onboarding from 2–3 weeks to minutes, but only because we built the right guardrails and transparency into the workflow so operators trusted the automation. Building repair queues and audit trails that give on-call engineers the context they need to act under pressure is a UX problem I find genuinely compelling, and one I have direct experience with from designing developer-facing platforms where the cost of confusion is measured in downtime.
**Selected Relevant Experience**
- **ICE Platform Scaling (Intuit):** Scaled throughput from 6K to 50K TPS via rSocket migration supporting ~1.5M concurrent connections at sub-25ms TP99; achieved 275% YoY growth in ICE engagements to 675M+ in FY23 across five major product lines.
- **MSaaS Drift Detection Program (Intuit):** Wrote Java JAR library to scan Git repositories for configuration drift across distributed microservices; built remediation roadmap using OpenRewrite — directly analogous to failure attribution and automated repair workflows.
- **ICE Self-Service Platform (Intuit):** Delivered DevPortal, GitOps config, and ICE Playground, reducing developer onboarding from 2–3 weeks to minutes in pre-prod and under 24 hours for production, while mitigating $1M+ in projected opex growth.
- **OpenClaw Multi-Agent Orchestration (StreamIO AI):** Designed and implemented multi-agent gateway protocol with subagent delegation, profile management, and session switching — foundational architecture for agentic workflow software at scale.
- **aeval Evaluation Platform:** Built automated safety gates, CI/CD regression detection, and statistical rigor (bootstrap confidence intervals, Welch's t-test, Cohen's d) into an AI evaluation pipeline — operationalizing the principle that automation must be measurably trustworthy before it acts.
- **Splunk Search Orchestration (Splunk Inc.):** Owned Go microservices for Search Service and Search Catalog; delivered Scheduler Service end-to-end in approximately four months; led query performance optimization achieving up to 10x improvements for enterprise beta customers.
- **Splunk Logging-as-a-Service (Kaiser Permanente):** Led enterprise rollout of a logging platform handling 1.7 TB daily volume across 200+ internal enterprise customers — operational scale with real SLA consequences.
**Closing**
NVIDIA's mission to build AI infrastructure that self-heals is not just an engineering ambition — it is the prerequisite for AI at the scale the world is moving toward. I want to be the product leader who defines how that self-healing system earns the trust of the operators who depend on it. I bring 12+ years of platform and infrastructure product leadership, hands-on distributed systems and agentic AI development, and a research foundation that keeps me grounded in the technical realities underneath the roadmap. I would welcome the opportunity to discuss how my background maps to what you're building.
Respectfully,
**O. Felix Amoruwa**
famoruwa@berkeley.edu | 909-731-9011 | felixamoruwa.info