jobsearch v0.0.1

← thinkingmachines / Product Manager - Deployment

cover_letter / art_QQx73fuKZe4

role
thinkingmachines / Product Manager - Deployment
model
anthropic/claude-sonnet-4.6
created
2026-09-16T19:32

↓ Download .docx

Cover letter

Dear Thinking Machines Hiring Team, Thinking Machines is building something I find genuinely compelling: frontier models paired with the infrastructure to make them useful in the real world, under the thesis that the future worth building is human. That framing resonates with me directly — I have spent the last several years not just building AI products, but operating them: running inference pipelines under load, debugging latency regressions at 2am, and learning firsthand that the gap between a trained model and a reliable production endpoint is where most of the hard work actually lives. **Technical and ML Foundation** My ML background starts at the infrastructure layer. In 2004 I hand-coded backpropagation through time in C++ for protein structure prediction — work that became a NeurIPS 2014 accepted paper and, in 2026, a full rewrite spanning 413 parameters to 8B (a 19-million-fold scale increase), now running across five architectures (feedforward, GRU, Transformer, ESM-2, multi-task) with MLflow experiment tracking, Optuna HPO, and FastAPI serving in a six-container Docker stack. That arc — from hand-rolled BPTT to modern serving infrastructure — gives me a concrete mental model for what happens at every layer between training and inference. More recently I built a post-training RL workbench that covers the full RLHF/DPO pipeline: a Reward Lab for designing and A/B testing reward functions across GSM8K, MATH, HumanEval, and UltraFeedback; a Playground running real TRL-powered GRPO/DPO training with live SSE metric streaming on Apple Silicon (MPS) and CUDA; and an Arena for head-to-head framework benchmarking across TRL, VeRL, OpenRLHF, and NeMo RL with GPU passthrough in Docker containers. I implemented 12 RL algorithms (PPO, GRPO, DAPO, REINFORCE, REINFORCE++, RLOO, DPO, SimPO, IPO, KTO, ORPO, SPPO) with standardized throughput, memory, and convergence benchmarking across frameworks. This is exactly the kind of work that informs deployment decisions — understanding how a training framework's throughput characteristics translate into serving cost and latency tradeoffs is not something you can reason about from a product brief alone. On the infrastructure side, at Intuit I scaled the ICE platform from 6K to 50K TPS via an rSocket migration supporting approximately 1.5M concurrent connections at sub-25ms TP99, achieving 275% YoY growth to 675M+ engagements in FY23 across QuickBooks, TurboTax, Mint, Mailchimp, and Credit Karma. I also built the aeval model evaluation platform — a local-first system with FastAPI orchestration, TimescaleDB, Redis job queuing, and automated safety gates — which gave me direct experience designing the observability and regression-detection infrastructure that production ML systems require. **Why This Role** Deployment at Thinking Machines is still being defined — what production-ready means for a fine-tuned checkpoint, which serving paths to support, how much control users get over performance and cost tradeoffs. That is precisely the kind of 0-to-1 infrastructure product problem I have navigated before, and I want to do it at a company working at the frontier. What excites me specifically about this role is the scope: owning the path from a Tinker-trained checkpoint to a served, reliable, cost-effective endpoint — covering autoscaling behavior, rollback safety, latency budgets, SLA definitions, and the API/SDK surfaces researchers and external users depend on. The JD's framing of "do whatever work makes deployment succeed — reviewing a serving config, joining an incident retro, inspecting latency data" matches how I actually work. I do not manage deployment from a distance; I have written the rollout plans, debugged the serving configs, and been in the incident retros. **Selected Prior Experience** - Scaled ICE platform throughput from 6K to 50K TPS via rSocket migration supporting ~1.5M concurrent connections at sub-25ms TP99; grew engagements to 675M+ in FY23 across five Intuit product lines. - Built RL post-training workbench benchmarking GRPO/DPO across TRL, VeRL, OpenRLHF, and NeMo RL with GPU Docker passthrough — standardized throughput, memory, and convergence metrics across 12 RL algorithms. - Built aeval evaluation platform with FastAPI orchestrator, TimescaleDB, Redis job queue, bootstrap confidence intervals, Welch's t-test, and automated CI/CD safety gates with regression detection. - Delivered ICE Self-Service platform (DevPortal, GitOps config, ICE Playground), reducing developer onboarding from 2–3 weeks to minutes in pre-prod and under 24 hours for production, while mitigating $1M+ in projected opex growth. - Engineered real-time HLS livestreaming pipeline with multi-stream canvas compositing (up to 9 concurrent sources at 30fps), FFmpeg WebM-to-MPEG-TS transcoding, and WebSocket layer — production infrastructure built and operated as a solo technical founder. - Architected RAG retrieval pipeline (ChromaDB) with multi-provider LLM orchestration (Claude, GPT-4, Gemini), fallback routing, structured-output validation, and token-budget optimization for Fintellect AI. - Led Mailchimp GCP-to-AWS migration for MSaaS, delivering Golang template, MySQL persistence integration, and updated DevPortal documentation to meet production deadline. **Closing** Thinking Machines' mission — AI that extends human will and judgment — requires that the models actually reach users reliably, predictably, and at cost. Deployment is not the last mile; it is the product. I have spent twelve years building toward the intersection of ML depth and infrastructure scale, and I want to bring that to Thinking Machines at a moment when the deployment model is still being written. I would welcome the opportunity to discuss how my background maps to what you are building. Sincerely, **O. Felix Amoruwa** famoruwa@berkeley.edu | 909-731-9011 | felixamoruwa.info