← crusoe / Staff Product Manager, AI Infrastructure (Storage)
cover_letter / art_NZD3aL3vWi8
role
model
anthropic/claude-sonnet-4.6
created
2026-08-31T19:05
Cover letter
Dear Crusoe Hiring Team,
Crusoe is doing something genuinely rare: building AI infrastructure from the ground up, owning every layer from power generation to compute, so that the most demanding AI workloads in the world have a foundation that can actually keep up with them. That vertical integration — and the conviction that energy and intelligence are the defining industrial challenge of this era — is exactly the kind of infrastructure-first thinking I want to work on. My own path from hand-coding backpropagation through time in C++ at UC Berkeley in 2004 to building production RL post-training workbenches that benchmark GRPO and DPO across TRL, VeRL, OpenRLHF, and NeMo RL today has kept me close to the infrastructure that makes AI work, not just the models on top of it.
---
**Technical and Product Foundation**
My most directly relevant infrastructure experience comes from Intuit, where I owned the developer platform and framework infrastructure serving QuickBooks, TurboTax, Mint, Mailchimp, and Credit Karma at scale. The work was fundamentally about keeping compute fed and developers unblocked — the same problem Crusoe's storage stack solves for GPU clusters. I scaled the ICE platform from 6K to 50K transactions per second via an rSocket migration supporting approximately 1.5 million concurrent connections at sub-25ms TP99, and grew platform engagements 275% year-over-year to 675M+ in FY23. I also led a GCP-to-AWS migration for Mailchimp's MSaaS workloads, delivering the Golang service template, MySQL persistence integration, and updated DevPortal documentation against a hard production deadline — the kind of cross-cloud infrastructure transition that maps directly to the storage tiering and lifecycle decisions Crusoe's customers face.
At Kaiser Permanente, I built and operated Splunk Logging-as-a-Service at 1.7 TB daily ingest volume across 200+ internal enterprise customers, and introduced Redis and XC10 caching across the enterprise to address scalability, fault tolerance, and data redundancy — problems that sit at the heart of durable, highly available storage design. At Splunk, I owned Search Service (Go microservices) and Search Catalog (PostgreSQL metadata service), and led a query performance optimization initiative that achieved up to 10x improvements for a beta Fortune 500 customer — work that required understanding data access patterns, indexing tradeoffs, and the economics of storage-backed search at scale.
On the AI research side, my NeurIPS 2014 paper on neural networks for protein structure prediction and my recent RL Workbench project — which benchmarks 12 algorithms across four frameworks with GPU Docker passthrough — reflect a habit of building rigorous, instrumented systems rather than one-off prototypes. My aeval evaluation platform (FastAPI, TimescaleDB, Redis, Ollama) and the BRAIN protein structure prediction platform (PyTorch, MLflow, Docker orchestration across 6 containers, 823 automated tests) demonstrate that I build production-grade infrastructure, not demos.
---
**Why This Role**
Storage is the unglamorous bottleneck that determines whether a training cluster runs at 90% GPU utilization or 60%. Crusoe's position — owning the full stack from electrons to tokens — means the storage product team has both the authority and the obligation to make decisions that most cloud vendors can only approximate through vendor negotiations. The scope of this role (block, file, and object across tiering, versioning, backup/restore, and lifecycle management) is exactly the kind of multi-dimensional infrastructure problem I find most interesting: technically deep, economically constrained, and directly tied to whether customers can ship.
I'm particularly drawn to the customer discovery mandate in this role. At Intuit, I conducted an enterprise-wide Service Language Assessment across nine languages, synthesizing usage telemetry and developer interviews into strategic recommendations presented to the CTO. I built Asterias, a declarative asset lifecycle management platform with a GraphQL API, directly from developer pain-point analysis using SQL and BigQuery. That pattern — close customer contact, data-grounded prioritization, and clear executive communication — is how I work, and it maps directly to what this role requires.
---
**Selected Relevant Experience**
- Scaled ICE platform throughput from 6K to 50K TPS via rSocket migration, supporting ~1.5M concurrent connections at sub-25ms TP99; grew platform to 675M+ engagements in FY23 across five major product lines.
- Built and operated Splunk Logging-as-a-Service at Kaiser Permanente: 1.7 TB daily ingest, 200+ enterprise customers, Redis/XC10 caching layer for fault tolerance and data redundancy.
- Led Mailchimp GCP-to-AWS MSaaS migration, delivering Golang service template, MySQL persistence, and DevPortal documentation against production deadline.
- Delivered ICE Self-Service DevPortal and GitOps configuration platform, reducing developer onboarding from 2–3 weeks to minutes in pre-prod and under 24 hours for production, while mitigating $1M+ in projected opex growth.
- Owned Search Service (Go microservices) and Search Catalog (PostgreSQL) at Splunk; led query performance initiative achieving up to 10x improvement for beta enterprise customer.
- Initiated MSaaS Drift Detection and Resolution program: authored Java JAR library to scan Git repos for configuration drift, partnered with Design on DevPortal UI, and built remediation roadmap using OpenRewrite.
- Conducted enterprise-wide Service Language Assessment across 9 languages, synthesizing usage data and developer feedback into strategic investment recommendations presented to CTO.
- Built production AI infrastructure across multiple projects: FastAPI + TimescaleDB + Redis evaluation platform (aeval), 6-container Docker orchestration for ML serving (BRAIN), and GPU Docker passthrough for multi-framework RL benchmarking.
---
Crusoe's mission — accelerating the abundance of energy and intelligence — is not a tagline; it is a genuine infrastructure thesis. The storage layer is where that thesis either holds or breaks under load. I want to own that problem: talk to the engineers whose training jobs are stalling on I/O, understand the economics of hot versus cold tiering at Crusoe's scale, and build a roadmap that makes the storage platform a reason customers choose Crusoe, not just a table-stakes requirement. I would welcome the opportunity to discuss how my background maps to what you are building.
Sincerely,
**O. Felix Amoruwa**
famoruwa@berkeley.edu · 909-731-9011 · felixamoruwa.info