jobsearch v0.0.1

← coreweave / Staff Product Manager, Insights

cover_letter / art_hzfwyX2zVCY

role
coreweave / Staff Product Manager, Insights
model
anthropic/claude-sonnet-4.6
created
2026-05-27T21:46

↓ Download .docx

Cover letter

Dear CoreWeave Hiring Team, CoreWeave has built something genuinely rare: a cloud platform purpose-engineered for the workloads that matter most right now — GPU-dense AI training, inference, and the infrastructure that frontier labs depend on. What draws me to this role is not abstract admiration for the company's trajectory; it is that I have been a CoreWeave customer in practice. My RL post-training workbench runs GRPO, DPO, and PPO across TRL, VeRL, OpenRLHF, and NeMo RL with GPU passthrough in Docker containers — exactly the kind of compute-intensive, observability-hungry workload your Insights team is building for. I know firsthand what it means to stare at raw telemetry from a GPU cluster and wish the platform would surface the signal instead of making you hunt for it. ## Technical and Observability Foundation My observability background runs deep and spans both the infrastructure and product dimensions this role requires. At Splunk, I owned Search Service (Go microservices), Search Catalog (PostgreSQL metadata), and Splunk Processing Language (SPL/SPL2) for Splunk Cloud Services. I built product roadmaps, PRDs, and acceptance criteria for the full search and scheduling stack, delivered the Scheduler Service end-to-end in roughly four months, and led a query performance optimization initiative that achieved up to 10x improvements for a beta enterprise customer. That work required translating raw log and metric pipelines into experiences that operators could actually act on — the same translation problem the Insights team solves every day. At Kaiser Permanente, I led the enterprise rollout of Splunk Logging-as-a-Service at 1.7 TB daily volume across 200+ internal customers, and built caching infrastructure using Redis and XC10 to address scalability and fault-tolerance at scale. I also led demand forecasting and capacity planning for the SOA product suite across multiple datacenters — work that required understanding not just what the telemetry said, but what it meant for cost and reliability decisions downstream. At Intuit, I scaled the ICE platform to 675M+ engagements in FY23 and drove throughput from 6K to 50K TPS via rSocket migration supporting approximately 1.5M concurrent connections at sub-25ms TP99. I worked directly with SQL and BigQuery telemetry to prioritize developer pain points across roughly 20 mobile apps and 30+ product SKUs. That kind of data-driven prioritization — using usage signals to decide what to build next — is exactly the muscle the Insights PM role requires. On the AI-powered insights side, I built aeval, a local-first model evaluation platform with a FastAPI orchestrator, TimescaleDB for time-series metrics, a Redis job queue, and an Ollama backend. The platform includes bootstrap confidence intervals, Welch's t-test, Cohen's d effect size, and automated safety gates integrated into CI/CD — statistical rigor applied to surfacing signal from model evaluation runs. I also architected a RAG retrieval pipeline with ChromaDB, multi-provider LLM orchestration across Claude, GPT-4, and Gemini with fallback routing, and structured output validation. These are not adjacent skills; they are directly applicable to building AI-powered insights that surface high-signal, low-noise information to customers operating complex GPU workloads. ## Why This Role The Insights PM role sits precisely at the intersection where my career has converged: infrastructure-scale telemetry, developer-facing platform products, and AI-powered analysis. The specific challenge of translating raw metrics, logs, and events into proactive, actionable insights for customers running AI workloads is one I have approached from multiple angles — as a platform PM at Splunk and Intuit, as a builder of evaluation and observability tooling, and as an end user of GPU cloud infrastructure running RL training jobs. What excites me most about this role is the cost optimization and workload efficiency domain. GPU compute is expensive, and customers running training and inference workloads on CoreWeave's Blackwell clusters need to understand not just that something went wrong, but why, and what to do about it. Building Grafana-based experiences that surface those signals proactively — rather than requiring customers to write their own PromQL queries after the fact — is a product problem with real, measurable customer impact. I want to own that problem. ## Selected Relevant Experience - **Splunk Search & Observability (2019–2021):** Owned Search Service (Go microservices), Search Catalog, and SPL/SPL2; delivered Scheduler Service end-to-end in ~4 months; led query performance optimization achieving up to 10x improvements for enterprise beta customer. - **Splunk Logging-as-a-Service at Kaiser (2012–2018):** Led enterprise rollout at 1.7 TB daily volume, 200+ internal customers; built Redis/XC10 caching layer for scalability and fault tolerance; led capacity planning across multiple datacenters. - **ICE Platform at Intuit (2021–2024):** Scaled to 675M+ engagements, 50K TPS, sub-25ms TP99; used BigQuery and SQL telemetry to prioritize across 20 mobile apps and 30+ SKUs; delivered $480K/month in additional invoicing via ICE Presence. - **aeval — AI Model Evaluation Platform (2025–2026):** Built local-first evaluation platform with FastAPI, TimescaleDB, Redis, and Ollama; implemented statistical rigor (bootstrap CI, Welch's t-test, Cohen's d) and automated CI/CD safety gates for regression detection. - **RL Workbench — Post-Training RL Platform (2026):** Built 3-phase workbench covering RLHF/DPO pipeline with live SSE metric streaming, GPU Docker passthrough, and head-to-head framework benchmarking across TRL, VeRL, OpenRLHF, and NeMo RL — direct experience as a CoreWeave-class GPU workload operator. - **Fintellect AI — RAG and LLM Orchestration (2024–Present):** Architected RAG pipeline with ChromaDB, multi-provider LLM orchestration with fallback routing, structured output validation, and token budget optimization — applicable to natural-language insight generation. - **Intuit ICE Self-Service Platform:** Reduced developer onboarding from 2–3 weeks to minutes; mitigated $1M+ in projected opex growth; built Asterias declarative asset lifecycle management platform with GraphQL API. ## Closing CoreWeave's mission — turning compute into capability for the pioneers building the next generation of AI — is one I want to contribute to directly. The Insights team's work is foundational to that mission: customers cannot operate complex GPU-powered systems with confidence if they cannot see clearly inside them. I have spent my career building the telemetry pipelines, developer platforms, and AI-powered analysis tools that make that visibility possible. I would welcome the opportunity to bring that experience to CoreWeave. Thank you for your consideration. --- **O. Felix Amoruwa** famoruwa@berkeley.edu | 909-731-9011 | felixamoruwa.info