← apptronik / Staff Product Manager, Applied AI
cover_letter / art_fpMM9l94Rkg
role
model
anthropic/claude-sonnet-4.6
created
2026-09-25T17:02
Cover letter
Dear Apptronik Hiring Team,
Apptronik is building something that matters: a humanoid robot designed to work alongside people in the environments where physical labor is hardest, most dangerous, or most repetitive. The Apollo platform sits at the intersection of physical AI, edge deployment, and real-world safety — exactly the convergence I have been working toward since hand-coding backpropagation through time in C++ at UC Berkeley in 2004 and, more recently, building a production RL post-training workbench that benchmarks GRPO, DPO, PPO, and nine other algorithms across TRL, VeRL, OpenRLHF, and NeMo RL on real hardware. The Staff Product Manager, Applied AI role is the natural next step in that arc.
**AI and ML Foundation**
My technical credibility in this space is grounded in hands-on work, not just product oversight. In 2025–2026 I built an RL post-training workbench from scratch covering the full RLHF/DPO pipeline: a Reward Lab for designing and A/B testing reward functions (RLVR, learned, hybrid) across GSM8K, MATH, HumanEval, and UltraFeedback; a Playground running real TRL-powered GRPO and DPO training with live SSE metric streaming on Apple Silicon (MPS) and CUDA; and an Arena for head-to-head framework benchmarking with GPU passthrough in Docker containers. I implemented 12 RL algorithms with algorithm-specific metric profiles, cross-tab workflow lineage tracking, and standardized throughput/memory/convergence benchmarking — the same class of evaluation rigor that gates model graduation in a physical AI stack.
On the evaluation side, I built aeval, a local-first model evaluation platform with five core eval types (factuality, reasoning, instruction-following, safety, code generation), adversarial safety testing with refusal detection, bootstrap confidence intervals, Welch's t-test, Cohen's d effect size, and CI/CD regression detection. Defining what "good" looks like before a model ships — and building the tooling to enforce that bar — is precisely what the JD describes as owning benchmarks and success criteria that gate real-world deployment.
Perhaps most directly relevant to Apptronik's evaluation challenges: I built AutoEval, an automated visual evaluation system for robot model training. By repurposing StreamIO's screen-capture and HLS streaming pipeline, I created a zero-integration architecture that scores model outputs — grasp poses, segmentation maps, bounding boxes — against natural-language evaluation rubrics using multimodal AI (Claude/GPT-4V), reducing evaluation cycles from 72 hours to approximately 4 minutes. That system captures from any visualization tool (RViz, Matplotlib, custom dashboards) via screen capture, which means it works with the tooling robotics teams already use rather than requiring instrumentation changes. This is the kind of practical, velocity-oriented solution the role calls for.
My NeurIPS 2014 publication on artificial neural networks for protein secondary structure prediction, and the 2026 rewrite of that system into a production PyTorch platform spanning 413 parameters to 8B (a 19-million-fold scale increase), establish that my engagement with neural architectures is longitudinal and substantive — not a recent pivot.
**Connecting the Arc**
The through-line in my career is translating complex technical systems into products that ship at scale and deliver measurable economic value — from scaling Intuit's ICE platform to 675M+ engagements and 50K TPS, to founding two AI-native companies and shipping production applications across iOS, macOS, Android, and web. The Apptronik role asks for someone who can own the physical AI model roadmap, sequence capability graduation against reliability and safety, and align data strategy with real-world deployment — that is the exact PM motion I have been executing, applied now to humanoid robotics.
**Why This Role**
What excites me most about this position is the data strategy ownership. Defining what to collect, why, and to what quality bar — then aligning data operations and engineering teams around that strategy — is where product management creates the most leverage in a model-driven system. Apollo's deployment in manufacturing and logistics means the data collection problem is tightly coupled to physical safety constraints (ISO 10218, ISO/TS 15066), and I find that constraint-driven design more interesting than unconstrained greenfield work. The requirement to translate customer use cases in industrial environments into model capability requirements and acceptance criteria is also a direct match to the PRD and acceptance-criteria work I did at Splunk for Search Orchestration and at Intuit for developer platform infrastructure.
**Selected Relevant Experience**
- **RL Workbench (2026):** Implemented 12 RL algorithms (PPO, GRPO, DAPO, REINFORCE, REINFORCE++, RLOO, DPO, SimPO, IPO, KTO, ORPO, SPPO) with standardized throughput/memory/convergence benchmarking across TRL, VeRL, OpenRLHF, and NeMo RL; built GPU Docker passthrough for framework-level Arena benchmarking.
- **AutoEval — Automated Visual Evaluation for Robot Model Training (2025):** Reduced robot model evaluation cycles from 72 hours to ~4 minutes using multimodal AI scoring of grasp poses, segmentation maps, and bounding boxes against natural-language rubrics; zero-integration capture from RViz, Matplotlib, and custom dashboards.
- **aeval — AI Model Evaluation Platform (2025–2026):** Defined and enforced model quality gates via bootstrap confidence intervals, Welch's t-test, Cohen's d, saturation detection, and CI/CD regression detection with automated safety gates.
- **Intuit ICE Platform (2021–2024):** Achieved 275% YoY growth in ICE engagements, scaling to 675M+ in FY23; scaled throughput from 6K to 50K TPS via rSocket migration supporting ~1.5M concurrent connections with sub-25ms TP99 — demonstrating the ability to own and execute a platform roadmap at production scale.
- **Intuit Developer Frameworks:** Delivered ICE Self-Service platform (DevPortal, GitOps config, ICE Playground), reducing developer onboarding from 2–3 weeks to minutes; extended Java and Python SDK Starter Kits with scaffolding, build configurations, and CI/CD integration.
- **Splunk Search Orchestration (2019–2021):** Owned Search Service (Go microservices), Search Catalog (PostgreSQL metadata service), and SPL/SPL2; delivered Scheduler Service end-to-end in ~4 months and achieved up to 10x query performance improvements for a Fortune 500 beta customer.
- **BRAIN Protein Structure Prediction Platform (UC Berkeley 2004; rewritten 2026):** NeurIPS-published; 2026 rewrite spans 413 to 8B parameters across five neural architectures (feedforward, GRU, Transformer, ESM-2, multi-task) with MLflow, Optuna HPO, and FastAPI serving — 823 automated tests, Docker orchestration across 6 containers.
**Closing**
Apptronik's mission — robots that support humanity in every facet of life, starting with the industries where the work is hardest — is one I want to contribute to directly. The problems of capability graduation, data quality strategy, and real-world evaluation gating are not abstract to me; I have built the tooling, written the acceptance criteria, and shipped the systems. I would welcome the opportunity to bring that foundation to Apollo.
Thank you for your consideration.
O. Felix Amoruwa
famoruwa@berkeley.edu | 909-731-9011 | felixamoruwa.info