jobsearch v0.0.1

← cerebrassystems / AI Models, Product Manager

cover_letter / art_yuuVn1myfjI

role
cerebrassystems / AI Models, Product Manager
model
anthropic/claude-sonnet-4.6
created
2026-05-22T15:38

↓ Download .docx

Cover letter

Dear Cerebras Systems Hiring Team, Cerebras is doing something structurally different from every other compute company: instead of asking developers to orchestrate fleets of GPUs, you've collapsed that complexity into a single wafer-scale device that delivers inference speeds an order of magnitude faster than GPU-based cloud services. That architectural bet is now paying off — the OpenAI partnership and the CS-3's 4 trillion transistors represent a genuine inflection point for what production AI inference can look like. My path into this work started in 2004, hand-coding backpropagation through time in C++ for a protein structure prediction system that eventually became a NeurIPS 2014 paper — and it has continued through building RL post-training infrastructure, multi-provider LLM orchestration pipelines, and a model evaluation platform from scratch. The AI Models PM role at Cerebras is the intersection of every thread I've been pulling on. **Technical Foundation** The most directly relevant project I can point to is my RL Workbench, a three-phase post-training platform I built to benchmark GRPO, DPO, PPO, DAPO, REINFORCE++, RLOO, SimPO, IPO, KTO, ORPO, and SPPO — twelve algorithms in total — across TRL, VeRL, OpenRLHF, and NeMo RL. The Reward Lab phase lets me design and A/B test reward functions (RLVR, learned, and hybrid) across GSM8K, MATH, HumanEval, and UltraFeedback. The Playground runs real TRL-powered training with live SSE metric streaming on Apple Silicon MPS or CUDA. The Arena benchmarks frameworks head-to-head with GPU passthrough in Docker containers, producing standardized throughput, memory, and convergence comparisons. This is the kind of systematic evaluation infrastructure that maps directly to what the JD describes: defining quality standards across a model catalog and designing benchmarks that prove production-grade performance. On the evaluation side, I built aeval — a local-first model evaluation platform covering factuality, reasoning, instruction-following, safety, and code generation, with adversarial safety testing, refusal detection, and data contamination detection via SHA-256 hashing. Statistical rigor was a design constraint from the start: bootstrap confidence intervals, Welch's t-test, Cohen's d effect size, and saturation detection. The stack is FastAPI orchestrator, TimescaleDB, Redis job queue, Next.js dashboard, and Ollama for local model serving. CI/CD integration includes regression detection and automated safety gates. Writing model quality evaluations and system prompt harnesses is listed as a preferred qualification in the JD — this is work I've done end-to-end. For multi-provider LLM orchestration, my Fintellect AI platform routes across Claude, GPT-4, and Gemini with fallback logic, structured output validation, and token budget optimization. My OpenClaw framework at StreamIO implements a multi-agent gateway protocol with subagent delegation, profile management, and session switching — coordinating agents across distinct task domains. These are not prototype integrations; they are production systems with real users. **Why This Role** My background spans the full stack the JD requires: hands-on ML research (NeurIPS, BRAIN protein structure platform with five neural architectures from feedforward to ESM-2), developer platform PM at scale (Intuit's ICE platform reaching 675M+ engagements at 50K TPS), and 0-to-1 product execution as a founder. The Cerebras AI Models PM role asks for someone who can own the model roadmap, define evaluation frameworks, drive go-to-market launches, and make technical tradeoffs on quantization and speculative decoding — that combination is exactly what I've been building toward. What specifically excites me about this role is the model portfolio strategy layer. Deciding which frontier and open-source models ship on Cerebras Inference — based on market demand, research trends, and strategic fit — is a problem that requires both deep ML fluency and product judgment about developer adoption curves. The day-0 launch partnerships with model labs, and the community relationships with open-source maintainers, are areas where my experience bridging technical depth with go-to-market execution is directly applicable. The performance optimization decisions — quantization, speculative decoding, latency/throughput tradeoffs — are questions I've engaged with in building inference pipelines and benchmarking RL frameworks. **Selected Prior Experience** - Built RL post-training workbench implementing 12 algorithms (PPO, GRPO, DAPO, DPO, SimPO, IPO, KTO, ORPO, SPPO, REINFORCE, REINFORCE++, RLOO) with standardized throughput/memory/convergence benchmarking across TRL, VeRL, OpenRLHF, and NeMo RL — directly applicable to model quality and performance evaluation at Cerebras. - Built aeval evaluation platform with 5 core eval types, adversarial safety testing, bootstrap confidence intervals, Welch's t-test, Cohen's d, and CI/CD regression detection — applicable to defining and enforcing quality standards across Cerebras' model catalog. - Architected multi-provider LLM orchestration pipeline (Claude, GPT-4, Gemini) with fallback routing, structured output validation, and token budget optimization at Fintellect AI — applicable to understanding inference tradeoffs customers face. - Delivered ICE Self-Service platform at Intuit, reducing developer onboarding from 2–3 weeks to minutes; scaled ICE to 675M+ engagements (275% YoY growth) and 50K TPS via rSocket migration — applicable to go-to-market execution and developer adoption strategy. - Extended Java and Python SDK Starter Kits with scaffolding, build configurations, testing frameworks, and CI/CD integration, enabling developers to reach production-ready microservices in minutes — applicable to creating developer-facing technical content and documentation. - NeurIPS 2014 accepted paper on artificial neural networks for protein secondary structure prediction; original 2004 C++ system with hand-coded BPTT rewritten in 2026 to span 413 parameters to 8B (19M-fold scale increase) — establishes foundational ML research credibility. - Led Splunk Search Orchestration product (Go microservices, PostgreSQL, SPL/SPL2), delivered Scheduler Service end-to-end in ~4 months, and achieved up to 10x query performance improvements for a beta enterprise customer — applicable to cross-functional launch execution and performance benchmarking. **Closing** Cerebras' thesis is that the constraint isn't the model — it's the compute substrate. By removing that constraint, you're enabling a class of agentic and real-time applications that simply weren't feasible on GPU-based inference. The AI Models PM role is the function that translates that hardware advantage into a model portfolio developers actually adopt and build on. I've spent the last two years building the evaluation infrastructure, orchestration frameworks, and RL post-training tooling that would let me contribute to that mission from day one. I'd welcome the opportunity to discuss how my background maps to Cerebras' model roadmap priorities. Sincerely, **O. Felix Amoruwa** famoruwa@berkeley.edu | 909-731-9011 | felixamoruwa.info