← roblox / Senior Product Manager, Compute Platform
tailored_resume_v2 / art_EMCcKK37-rE
role
model
anthropic/claude-sonnet-4.6
created
2026-06-15T19:25
↓ Download .docx ↓ Download .pdf PDF requires LibreOffice installed
What changed for roblox
| change | why it matters |
|---|---|
| Summary rewritten to lead with compute platform scale (675M+ engagements, 50K TPS) and GPU training framework benchmarking | JD's first hard requirement is 7+ years PM experience on compute infrastructure or distributed systems at scale; GPU infrastructure is the role's core |
| Intuit role moved to lead Experience section and first bullet anchors platform scale metrics | Intuit is the strongest proof of production-grade compute platform PM at scale — directly maps to Roblox's reliability and cost-to-serve requirements |
| Intuit role title reframed to 'Developer Frameworks & Platform Infrastructure' (already accurate) with bullets reordered to lead with scale, then onboarding, then cloud migration | JD prioritizes compute reliability, developer experience, and on-prem/cloud infrastructure in that order |
| Intuit GCP-to-AWS migration bullet explicitly labeled 'hands-on cloud infrastructure delivery across on-prem and cloud' | JD preferred qual: experience building compute infrastructure on AWS/GCP; this is the strongest factual evidence |
| Intuit Drift Detection bullet reframed as 'fleet-wide health and configuration management at scale' | JD explicitly calls out fleet-wide health and performance as a core ownership area |
| Streamio AI role title reframed to 'Compute & Agentic AI Platform' and OpenClaw bullet leads | JD preferred qual: experience building agentic systems for compute or infrastructure; OpenClaw is the strongest proof point |
| Splunk role title reframed to include 'Distributed Systems' | JD hard requirement: distributed systems at scale; Go microservices ownership is the strongest Splunk proof point |
| Kaiser role title reframed to 'Infrastructure & Platform Services'; observability bullet leads | Fleet-wide health/observability and capacity planning across datacenters maps to JD's reliability and fleet management requirements |
| RL Workbench project moved to lead the Projects section | GPU Docker passthrough, CUDA/MPS hardware benchmarking, and framework throughput/memory/convergence metrics are the strongest proof of GPU infrastructure depth — JD's most differentiated technical requirement |
| RL Workbench project title reframed to 'GPU Post-Training Compute Platform' | Signals GPU infrastructure ownership rather than just ML research, matching JD's compute platform framing |
| BRAIN project's C++ BPTT bullet explicitly noted as 'kernel-level and hardware-proximate ML engineering background' | JD preferred qual: kernel-level experience or familiarity with custom kernel drivers |
| Fintellect AI role removed from Experience section | Lowest relevance score (2) among founder roles; RAG/LLM orchestration covered by RL Workbench and aeval projects; removing creates space for stronger compute infrastructure content within 2-page target |
| Intuit ICE engagements bullet annotated with Roblox DAU comparison context | Helps hiring manager immediately map candidate's scale experience to Roblox's 88M DAU platform engineering challenges |
| RL Workbench second bullet annotated with 'directly applicable to Roblox's GPU fleet optimization and cost-to-serve' | JD explicitly calls out cost-to-serve optimization as a success metric; throughput/memory benchmarking across frameworks is the strongest factual proof |
JD analysis (19 key phrases)
Key phrases: Compute PlatformGPU infrastructurefleet-wide health and performanceManaged Kubernetesdistributed systems at scaletraining and inference for frontier modelscost-to-servemean-time-to-detection and recoveryCompute Platform Reliabilitytopology-aware placementproduction-ready AI computedriver and firmware managementFleet APIson-prem and cloudplatform primitivesdeveloper experienceagentic systemsGPU and AI acceleratorsCompute infrastructure
Hard requirements:
- 7+ years product management experience focused on compute infrastructure or distributed systems at scale
- Deep practical understanding of Kubernetes internals and control plane components
- Experience productizing custom Kubernetes Operators, Controllers, and CRDs
- Deep familiarity with GPU/accelerator architecture and scheduling challenges
- Familiarity with cloud-native service networking including microservices, CNIs, service mesh
- Built production-grade compute platforms with reliability and developer experience
- Ability to balance utilization, latency, cost, and time-to-ship tradeoffs
- Builder mindset with rapid iteration and AI-leveraged ideation
Preferred qualifications:
- Experience building compute infrastructure on AWS, GCP, Azure
- Background in AI model development, training, inference
- Kernel-level experience or familiarity with custom kernel drivers
- Experience building agentic systems for compute or infrastructure
Per-role mapping (10 roles scored)
| role | score | reframe angle | JD phrases that map |
|---|---|---|---|
| Intuit — Staff Product Manager, Developer Frameworks & Platform Infrastructure | 5/5 | Production-grade compute/platform infrastructure PM at scale — reliability, developer experience, distributed systems, cloud migration | distributed systems at scale, production-ready, developer experience, on-prem and cloud, platform primitives, cost-to-serve, Compute Platform Reliability, Fleet APIs |
| Splunk — Senior Product Manager, Search Orchestration | 4/5 | Distributed systems PM — Go microservices, scheduling, query performance at scale | distributed systems at scale, mean-time-to-detection and recovery, topology-aware placement, Compute Platform Reliability, platform primitives |
| Kaiser Permanente — SOA Technical Product Manager | 3/5 | Infrastructure platform PM — capacity planning, fault tolerance, enterprise-scale observability | fleet-wide health and performance, Compute Platform Reliability, on-prem and cloud, mean-time-to-detection and recovery |
| Streamio AI — Founder & CEO | 4/5 | Agentic AI systems builder — multi-agent orchestration, real-time infrastructure, developer tooling | agentic systems, production-ready AI compute, developer experience, platform primitives |
| Fintellect AI — Founder & CEO | 2/5 | AI inference and multi-model orchestration at the application layer | training and inference for frontier models, agentic systems |
| IBM — Software Engineer, Business Intelligence Products | 2/5 | Enterprise software reliability and escalation resolution | mean-time-to-detection and recovery, Compute Platform Reliability |
| Bank of America Merrill Lynch — Tech MBA Summer Associate | 1/5 | Quantitative modeling and cost optimization | cost-to-serve |
| RL Workbench — Post-Training RL Platform | 5/5 | GPU compute infrastructure for AI training — framework benchmarking, throughput optimization, hardware-aware scheduling | GPU infrastructure, training and inference for frontier models, cost-to-serve, fleet-wide health and performance, production-ready AI compute |
| aeval — AI Model Evaluation Platform | 3/5 | Production AI infrastructure with reliability gates and observability | production-ready AI compute, Compute Platform Reliability, mean-time-to-detection and recovery |
| BRAIN — Protein Structure Prediction ML Platform | 3/5 | Full-stack ML platform engineering from training to serving, with deep hardware roots | training and inference for frontier models, production-ready AI compute, GPU infrastructure |
Tailored summary
Technical Product Leader with 12+ years building production-grade compute platforms and distributed systems at scale — from scaling Intuit's ICE platform to 675M+ engagements at 50K TPS to benchmarking GPU training frameworks (TRL, VeRL, OpenRLHF, NeMo RL) across CUDA/MPS hardware today. Deep experience driving Compute Platform Reliability, developer experience, and cloud infrastructure strategy (GCP, AWS) across on-prem and cloud. NeurIPS published AI researcher with hands-on background in AI training, inference, and agentic systems. CMU MS + MBA, UC Berkeley BS Engineering.