jobsearch v0.0.1

← coreweave / Staff Product Manager, Insights

cover_letter / art_eL_FCzKkeCY

role
coreweave / Staff Product Manager, Insights
model
anthropic/claude-sonnet-4.6
created
2026-05-27T21:43

↓ Download .docx

Cover letter

Dear CoreWeave Hiring Team, CoreWeave is building the essential infrastructure layer for the AI era — GPU-dense, purpose-built, and trusted by the frontier labs defining what's possible. That mission matters because the next generation of AI breakthroughs will be won or lost at the infrastructure layer, and the teams operating those systems need observability that matches the complexity of what they're running. My interest in this role is grounded in a specific experience: building my RL post-training workbench, I ran GRPO and DPO training jobs across TRL, VeRL, OpenRLHF, and NeMo RL on Apple Silicon (MPS) and CUDA, and discovered firsthand how opaque GPU workload behavior is without purpose-built instrumentation. Watching throughput, memory, and convergence metrics stream live during training made clear that the gap between raw telemetry and actionable insight is exactly where product work is most needed — and most underdone. **Technical and AI Foundation** My technical foundation spans the full stack relevant to this role. At Intuit, I owned the ICE platform — a developer infrastructure layer that scaled from 6K to 50K TPS via rSocket migration, supporting approximately 1.5M concurrent connections with sub-25ms TP99 latency, and reached 675M+ engagements in FY23 across QuickBooks, TurboTax, Mint, Mailchimp, and Credit Karma. I worked directly with telemetry and usage data in SQL and BigQuery to prioritize developer pain points across roughly 20 mobile apps and 30+ product SKUs. That experience — turning high-volume, multi-dimensional signals into prioritized product decisions — is precisely the muscle the Insights role requires. At Splunk, I owned Search Service (Go microservices), Search Catalog (PostgreSQL metadata service), and SPL/SPL2 — the query language that sits at the center of how operators make sense of log and metric data. I led a query performance optimization initiative that achieved up to 10x improvements in Splunk Cloud Services search for a beta enterprise customer, and I delivered the Scheduler Service end-to-end in approximately four months. That work gave me a practitioner's understanding of how observability pipelines are built, where they break under load, and what customers actually need when they're debugging production systems. On the AI side, I built aeval, a local-first model evaluation platform with FastAPI orchestration, TimescaleDB, Redis job queuing, and a Next.js dashboard — incorporating bootstrap confidence intervals, Welch's t-test, Cohen's d effect size, and automated safety gates. I also built the RL Workbench covering 12 algorithms (PPO, GRPO, DAPO, DPO, SimPO, and others) with live SSE metric streaming, cross-tab workflow lineage tracking, and standardized throughput/memory/convergence benchmarking across frameworks in GPU Docker containers. These are not adjacent projects — they are direct analogs to what CoreWeave's Insights team is building for customers running AI workloads at scale. **Why This Role** My arc from infrastructure engineering to platform product management to AI systems has consistently pointed toward the same problem: complex technical systems generate enormous signal, and the gap between raw data and human-actionable insight is where value is created or destroyed. The CoreWeave Insights role sits exactly at that intersection — infrastructure, observability, and UX — and the specific problem areas called out in the JD (cost optimization, workload efficiency, proactive signal surfacing) are problems I have encountered as both a builder and a PM. What excites me most about this role is the mandate to move observability from reactive to proactive — surfacing high-signal, low-noise information before customers have to go looking for it. The combination of Grafana-based experiences, natural-language analysis, and AI-powered alerting across GPU workloads is a genuinely hard product problem, and CoreWeave's position — with deep NVIDIA Blackwell partnerships and contracts with frontier AI labs — means the customer base is operating at the edge of what's possible. That's the environment where this kind of product work has the most leverage. **Selected Relevant Experience** - **ICE Platform Scale (Intuit):** Scaled throughput from 6K to 50K TPS via rSocket migration supporting ~1.5M concurrent connections with sub-25ms TP99; achieved 275% YoY growth in ICE engagements to 675M+ in FY23. - **Telemetry-Driven Prioritization (Intuit):** Used SQL and BigQuery to analyze usage data across ~20 mobile apps and 30+ product SKUs; built Asterias, a declarative asset lifecycle management platform with GraphQL API. - **Search Observability (Splunk):** Owned Search Service (Go microservices), Search Catalog (PostgreSQL), and SPL/SPL2; led query performance initiative achieving up to 10x improvements in SCS search for enterprise beta customer. - **RL Workbench — GPU Workload Benchmarking (2026):** Built 3-phase post-training platform with live SSE metric streaming, cross-tab workflow lineage, and standardized throughput/memory/convergence benchmarking across TRL, VeRL, OpenRLHF, and NeMo RL in GPU Docker containers. - **aeval — Model Evaluation Platform (2025–2026):** Built evaluation platform with FastAPI orchestration, TimescaleDB, Redis job queue, and statistical rigor (bootstrap CIs, Welch's t-test, Cohen's d); CI/CD integration with regression detection and automated safety gates. - **Developer Onboarding Observability (Intuit):** Delivered ICE Self-Service platform (DevPortal, GitOps config, ICE Playground), reducing developer onboarding from 2–3 weeks to minutes in pre-prod and under 24 hours for production, while mitigating $1M+ in projected opex growth. - **Logging-as-a-Service (Kaiser Permanente):** Led enterprise rollout of Splunk Logging-as-a-Service handling 1.7 TB daily volume across 200+ internal enterprise customers; built caching capability using Redis and XC10 for scalability and fault tolerance. **Closing** CoreWeave's mission — turning compute into capability for the pioneers building AI — depends on customers being able to operate complex GPU environments with confidence. The Insights team is the product surface that makes that confidence possible. I'd welcome the opportunity to discuss how my background in observability platforms, AI workload tooling, and high-scale infrastructure product management maps to what you're building. Thank you for your consideration. O. Felix Amoruwa famoruwa@berkeley.edu | 909-731-9011 | felixamoruwa.info