← reflectionai / Product Manager - OSS
cover_letter / art_kr-fgmtUYuE
Cover letter
Dear Reflection AI Hiring Team,
Reflection's mission — building open superintelligence and making it accessible to all — sits at the intersection of two things I care about deeply: rigorous AI research and the belief that open systems compound faster than closed ones. My path from hand-coding backpropagation through time in C++ at UC Berkeley in 2004, to publishing at NeurIPS, to building RL post-training workbenches that benchmark GRPO, DPO, PPO, and DAPO across TRL, VeRL, OpenRLHF, and NeMo RL today, has been a continuous bet on that same thesis. The PM role at Reflection is the role I would design for myself.
## Technical Foundation
My AI/ML work is not adjacent to my product work — it is the same work. In 2026 I built an RL post-training workbench covering the full RLHF pipeline end-to-end: a Reward Lab for designing and A/B testing reward functions (RLVR, learned, hybrid) across GSM8K, MATH, HumanEval, and UltraFeedback; a Playground running real TRL-powered GRPO and DPO training with live SSE metric streaming on Apple Silicon (MPS) and CUDA; and an Arena for head-to-head framework benchmarking across TRL, VeRL, OpenRLHF, and NeMo RL with GPU passthrough in Docker containers. I implemented 12 RL algorithms — PPO, GRPO, DAPO, REINFORCE, REINFORCE++, RLOO, DPO, SimPO, IPO, KTO, ORPO, SPPO — with algorithm-specific metric profiles and standardized throughput, memory, and convergence benchmarking across frameworks. That is not a survey project. It is the kind of hands-on fluency that lets me sit across from a research lead and have a real conversation about what to cut and why.
Before that, I built aeval, a local-first model evaluation platform with five core eval types (factuality, reasoning, instruction-following, safety, code generation), adversarial safety testing with refusal detection, and data contamination detection via SHA-256 hashing. Statistical rigor was non-negotiable: bootstrap confidence intervals, Welch's t-test, Cohen's d effect size, and saturation detection, with CI/CD integration for regression detection and automated safety gates. The stack — FastAPI orchestrator, TimescaleDB, Redis job queue, Next.js dashboard, Ollama — was chosen to keep the platform self-hostable and reproducible, which is the only honest way to do evals.
The longer arc: my original 2004 BRAIN system was a hand-coded neural network in C++ with custom BPTT for protein secondary structure prediction. The 2026 rewrite spans five architectures (feedforward, GRU, Transformer, ESM-2, multi-task), MLflow experiment tracking, Optuna HPO, and FastAPI serving — scaling from 413 to 8 billion parameters, a 19-million-fold increase. The NeurIPS 2014 paper on artificial neural networks for protein secondary structure prediction came from that lineage.
## Why This Role
I have spent the last three years watching the OSS AI ecosystem develop its own gravity — where downstream adoption, contribution health, and inference provider relationships matter far more than GitHub stars. The Reflection PM charter — owning model releases end-to-end, shaping what gets built with research, managing licensing and community, building ecosystem partnerships, and leading GTM without a layer of specialists to hand off to — is exactly the scope I want. I am not looking for a narrower role.
What excites me specifically: the combination of working directly with research leads to decide what gets cut and when, and then owning the open-source release through to community health and commercial reciprocity. Those two things are usually separated organizationally, and that separation is where OSS AI products lose coherence. Reflection has structured the role to keep them together, which is the right call.
## Selected Relevant Experience
- **RL Workbench (2026):** Implemented 12 RL algorithms with cross-tab workflow lineage tracking and standardized benchmarking across TRL, VeRL, OpenRLHF, and NeMo RL — directly applicable to evaluating and releasing open-weight model variants.
- **aeval (2025–2026):** Built a self-hostable model evaluation platform with adversarial safety testing, statistical rigor (bootstrap CIs, Welch's t-test, Cohen's d), and CI/CD regression detection — the kind of instrumentation that tells you whether a model release bet is working, not just whether it shipped.
- **Intuit — ICE Self-Service Platform:** Reduced developer onboarding from 2–3 weeks to minutes in pre-prod and under 24 hours for production; scaled ICE engagements 275% YoY to 675M+ in FY23; scaled throughput from 6K to 50K TPS via rSocket migration supporting ~1.5M concurrent connections at sub-25ms TP99. Developer-facing platform ownership at this scale requires the same instincts Reflection needs for inference provider and cloud partner relationships.
- **Intuit — Java and Python SDK Starter Kits:** Extended SDK scaffolding with build configurations, testing frameworks, and CI/CD integration — empowering developers to go from zero to production-ready microservice in minutes. This is the kind of developer experience work that drives downstream adoption in OSS ecosystems.
- **Splunk — Search Orchestration:** Owned Search Service (Go microservices), Search Catalog (PostgreSQL metadata service), and SPL/SPL2; delivered the Scheduler Service end-to-end in approximately four months; achieved up to 10x query performance improvements for a beta enterprise customer through benchmark-driven optimization. Spoke at Splunk .conf18 and .conf19 — public-facing credibility with technical audiences.
- **Fintellect AI — RAG and Multi-Provider LLM Orchestration:** Architected a RAG retrieval pipeline with ChromaDB vector store, multi-provider LLM orchestration (Claude, GPT-4, Gemini) with fallback routing, structured output validation, and token budget optimization — hands-on experience with the inference provider landscape Reflection's ecosystem partners operate in.
- **NeurIPS 2014:** Published paper on artificial neural networks for protein secondary structure prediction — establishes research credibility for working alongside AI researchers, not just alongside them organizationally.
## Closing
Reflection's bet is that open superintelligence, built rigorously and released accessibly, is the right foundation for what comes next. I share that bet, and I have been building toward the technical and product depth to contribute to it meaningfully — from the first BPTT loop in C++ to benchmarking GRPO across four frameworks today. I would welcome the opportunity to discuss how my background maps to what you are building.
---
**O. Felix Amoruwa**
famoruwa@berkeley.edu | 909-731-9011 | felixamoruwa.info