jobsearch v0.0.1

← hud / Product Manager, Platform

cover_letter / art_c7MA1ylnj3Y

role
hud / Product Manager, Platform
model
anthropic/claude-sonnet-4.6
created
2026-09-28T05:52

↓ Download .docx

Cover letter

Dear HUD Hiring Team, HUD is building the infrastructure layer that makes RL training data and evaluations reliable and tradeable — the kind of foundational work that determines whether frontier AI agents actually improve. That mission connects directly to work I have been doing for the past two years: building an RL post-training workbench that benchmarks GRPO, DPO, and ten other algorithms across TRL, VeRL, OpenRLHF, and NeMo RL, and shipping an evaluation platform (aeval) with statistical rigor — bootstrap confidence intervals, Welch's t-test, Cohen's d — designed to make model quality legible to practitioners who need to act on it. When I read about HUD's marketplace for RL environments and training data, I recognized the exact problem I have been building tooling around. --- **Technical and AI/ML Foundation** My engagement with RL and evaluation is hands-on and longitudinal. The RL Workbench I built in 2026 covers the full RLHF/DPO pipeline across three phases: a Reward Lab for designing and A/B testing reward functions (RLVR, learned, and hybrid) across GSM8K, MATH, HumanEval, and UltraFeedback; a Playground for real TRL-powered GRPO/DPO training with live SSE metric streaming on Apple Silicon (MPS) and CUDA; and an Arena for head-to-head framework benchmarking with GPU passthrough in Docker containers. Implementing 12 RL algorithms with algorithm-specific metric profiles and standardized throughput/memory/convergence benchmarking gave me a practitioner's understanding of what makes RL training data and evaluation environments useful — and what makes them frustrating to work with. The aeval platform extends that into structured evaluation: five core eval types, adversarial safety testing with refusal detection, data contamination detection via SHA-256 hashing, and CI/CD integration with regression detection and automated safety gates. Both projects were built to be used by researchers who need to compare outputs, not just run them — which is precisely the user HUD's marketplace serves. This work sits on top of a longer arc. I hand-coded backpropagation through time in C++ at UC Berkeley in 2004, published at NeurIPS 2014 on neural networks for protein secondary structure prediction, and rewrote that system in 2026 as a full PyTorch platform spanning 413 parameters to 8B — a 19-million-fold scale increase — with MLflow experiment tracking, Optuna HPO, and FastAPI serving. --- **Bridge** That technical depth informs how I approach product work: I can read a PRD, sit in an architecture review, and talk to a researcher about evaluation methodology without losing the thread — which matters when the product you are building is infrastructure for people who will immediately probe its assumptions. --- **Why This Role** The HUD marketplace has a two-sided structure — data vendors who create and submit environments, and research teams who discover and consume them — and the PM role is explicitly about making both sides work. That is the kind of problem I find most tractable: not a single user journey, but a system where the quality of one side's experience directly determines the value the other side receives. The specific work of improving vendor onboarding, quality feedback loops, and how research teams assess and access environments maps cleanly onto platform product work I have done, and onto the RL/eval domain I have been building in. --- **Selected Prior Experience** - **RL Workbench (2026):** Built a 3-phase post-training platform implementing 12 RL algorithms (PPO, GRPO, DAPO, DPO, SimPO, IPO, KTO, ORPO, SPPO, REINFORCE, REINFORCE++, RLOO) with cross-tab workflow lineage tracking and standardized benchmarking across TRL, VeRL, OpenRLHF, and NeMo RL — directly relevant to understanding what HUD's vendor and buyer users need from RL environment tooling. - **aeval — AI Model Evaluation Platform (2025–2026):** Shipped a local-first evaluation platform with statistical rigor (bootstrap CIs, Welch's t-test, Cohen's d), adversarial safety testing, and CI/CD regression gates — built to make evaluation results actionable for practitioners. - **Intuit — ICE Self-Service Platform:** Delivered a developer self-service platform (DevPortal, GitOps config, ICE Playground) that reduced onboarding from 2–3 weeks to minutes in pre-prod and under 24 hours for production, while mitigating $1M+ in projected opex growth — a direct analogue to improving vendor onboarding on HUD's marketplace. - **Intuit — Platform Scaling:** Achieved 275% YoY growth in ICE engagements, scaling to 675M+ in FY23; scaled throughput from 6K to 50K TPS via rSocket migration supporting ~1.5M concurrent connections with sub-25ms TP99 — evidence of owning a live platform product through growth and iteration. - **Intuit — Enterprise Service Language Assessment:** Conducted an enterprise-wide assessment across 9 languages, analyzing usage data and developer feedback to inform strategic investment decisions presented to the CTO — the kind of qualitative-plus-quantitative synthesis the HUD PM role requires. - **Splunk — Search Orchestration:** Owned three microservice product areas (Search Service in Go, Search Catalog in PostgreSQL, SPL/SPL2), delivered the Scheduler Service end-to-end in ~4 months, and led a query performance initiative achieving up to 10x improvements for a beta customer — experience turning complex technical requirements into shipped product. - **Vantage / Fintellect (Streamio AI, 2024–Present):** Led 0-to-1 product strategy, customer discovery, and go-to-market execution across two shipped products (iOS, macOS, web), including iterative refinement based on direct user interviews — the discovery-through-launch-and-iteration cycle the JD calls out explicitly. --- **Closing** HUD is building the data infrastructure layer that RL-trained frontier agents depend on. Getting that marketplace right — so vendors can submit high-quality environments and researchers can find and trust what they access — is consequential work. I have spent the last two years building in this exact technical domain, and the prior twelve building and scaling developer-facing platforms. I would welcome the chance to bring both to HUD. Thank you for your consideration. O. Felix Amoruwa famoruwa@berkeley.edu · 909-731-9011 · felixamoruwa.info