← cohere / Product Manager, Safety & Security
cover_letter / art_aZRX3HiE4MA
role
model
anthropic/claude-sonnet-4.6
created
2026-05-29T18:58
Cover letter
Dear Cohere Hiring Team,
Cohere's mission — scaling intelligence to serve humanity — sits at a genuinely hard intersection: frontier model capability and the discipline to deploy it responsibly. The North platform, in particular, represents exactly the kind of enterprise-grade agentic system where safety is not a feature to be bolted on but an architectural commitment that has to be built from the ground up. My path from hand-coding backpropagation through time in C++ at UC Berkeley in 2004, through a NeurIPS publication on neural networks for protein structure prediction, to building production RL post-training workbenches and AI evaluation platforms today has given me both the technical grounding and the product instincts this role demands.
**Technical and AI/ML Foundation**
My most directly relevant recent work is aeval, a local-first AI model evaluation platform I built from scratch in 2025–2026. aeval covers five core evaluation types — factuality, reasoning, instruction-following, safety, and code generation — with adversarial safety testing including refusal detection and data contamination detection via SHA-256 hashing. The statistical rigor layer includes bootstrap confidence intervals, Welch's t-test, Cohen's d effect size, and saturation detection, with CI/CD integration for regression detection and automated safety gates. This is the kind of evaluation infrastructure that sits directly upstream of the work Cohere's safety research teams are doing: defining what "safe" means quantitatively, detecting regressions before they reach customers, and building processes that scale.
Alongside aeval, I built a full RL post-training workbench covering the RLHF/DPO pipeline end-to-end — a Reward Lab for designing and A/B testing reward functions (RLVR, learned, hybrid) across GSM8K, MATH, HumanEval, and UltraFeedback; a Playground for live TRL-powered GRPO/DPO training with SSE metric streaming; and an Arena for head-to-head framework benchmarking across TRL, VeRL, OpenRLHF, and NeMo RL. Implementing 12 RL algorithms (PPO, GRPO, DAPO, REINFORCE, REINFORCE++, RLOO, DPO, SimPO, IPO, KTO, ORPO, SPPO) with standardized throughput, memory, and convergence benchmarking gave me a concrete understanding of how post-training choices shape model behavior — which is precisely the knowledge needed to ask the right questions when safety researchers surface unexpected findings from evaluations.
I also built OpenClaw, a multi-agent orchestration framework with a gateway protocol, subagent delegation, profile management, and session switching. Working through the design of a multi-agent system firsthand — where tool use, multi-step reasoning, and autonomous execution introduce failure modes that single-turn inference does not — gave me direct exposure to the safety challenges that are unique to agentic contexts, including the kinds of prompt injection and misuse patterns that North's safety roadmap will need to address.
**Why This Role**
The Safety Research PM role at Cohere is not a traditional PM seat, and that is what makes it compelling. The job description is explicit: you will spend as much time reading evaluations and engaging with researchers as writing PRDs. That framing matches how I have operated across my career — at Intuit, I worked directly with telemetry and usage data in BigQuery and SQL to surface developer pain points, and I built evaluation and benchmarking infrastructure rather than simply consuming outputs from other teams. The bridge between research findings and product-level guardrails is exactly the gap I want to close.
What specifically draws me to North is the agentic surface area. Autonomous agents executing multi-step workflows inside enterprise infrastructure introduce threat vectors — prompt injection, RAG poisoning, misuse through tool chaining — that are qualitatively different from those in single-turn chat products. My hands-on experience building and stress-testing multi-agent systems (OpenClaw, the RL workbench's framework benchmarking, AutoEval's zero-integration screen-capture pipeline) means I can engage with Cohere's safety researchers on these topics with genuine technical depth, not just product instinct.
**Selected Prior Experience**
- Built aeval, a production AI model evaluation platform with adversarial safety testing, refusal detection, data contamination detection, bootstrap confidence intervals, Welch's t-test, Cohen's d effect size, and automated safety gates integrated into CI/CD pipelines.
- Implemented 12 RL post-training algorithms across TRL, VeRL, OpenRLHF, and NeMo RL with standardized benchmarking for throughput, memory, and convergence — enabling rigorous, reproducible comparison of how training choices affect model behavior.
- Built OpenClaw multi-agent orchestration framework with gateway protocol and subagent delegation, gaining firsthand exposure to the safety failure modes introduced by tool use and autonomous multi-step execution.
- At Intuit, delivered ICE Self-Service platform reducing developer onboarding from 2–3 weeks to minutes, scaled throughput from 6K to 50K TPS, and achieved 675M+ engagements in FY23 — demonstrating the ability to own platform infrastructure at enterprise scale and translate technical complexity into product roadmaps.
- Initiated and led MSaaS Drift Detection and Resolution program at Intuit: wrote a Java JAR library to scan Git repositories for configuration drift and built a remediation roadmap — a direct analog to the kind of systematic safety regression detection this role requires.
- Architected RAG retrieval pipeline at Fintellect AI with ChromaDB vector store, multi-provider LLM orchestration (Claude, GPT-4, Gemini) with fallback routing, structured output validation, and token budget optimization — including hands-on exposure to RAG-specific failure modes.
- NeurIPS 2014 published researcher (neural networks for protein secondary structure prediction), with original C++ BPTT implementation dating to 2004 — establishing a research foundation that supports credible engagement with modeling and safety research teams.
**Closing**
Cohere's position — building frontier models for enterprise deployment while taking seriously the responsibility that comes with that — is where I want to direct the next chapter of my work. The Safety Research PM role offers the opportunity to do something that matters: ensure that as North's agentic surface area grows, the safety infrastructure grows with it, grounded in what the research is actually surfacing rather than in assumptions about what the risks might be. I would welcome the opportunity to discuss how my background maps to what Cohere needs.
Sincerely,
**O. Felix Amoruwa**
famoruwa@berkeley.edu | 909-731-9011 | felixamoruwa.info