← cohere / Product Manager, Safety & Security
tailored_resume_v2 / art_DHa3vcdL3xw
role
model
anthropic/claude-sonnet-4.6
created
2026-05-29T18:59
↓ Download .docx ↓ Download .pdf PDF requires LibreOffice installed
What changed for cohere
| change | why it matters |
|---|---|
| Section order flipped: Projects leads before Experience | JD explicitly values hands-on LLM evaluation, red-teaming, and safety benchmarking as preferred quals — aeval and RL Workbench are the strongest proof points and must appear first |
| Section retitled to 'AI Safety Research & Projects' | Mirrors JD's 'safety research' framing and signals alignment from the first glance |
| aeval project moved to lead position and reframed around adversarial safety testing, refusal detection, and automated safety gates | aeval is the single most relevant credential — it is literally a safety evaluation platform matching Cohere's core mandate |
| RL Workbench reframed as 'Post-Training RL & Safety Alignment Platform' with reward function A/B testing linked to safety reward modeling | JD requires understanding RLHF/alignment evaluation; reward function benchmarking maps directly |
| StreamIO/OpenClaw reframed as 'Multi-Agent Orchestration & Agentic Safety' project | JD explicitly calls out agentic AI systems, tool use, multi-step reasoning, and autonomous execution as a preferred qual area |
| Fintellect reframed around RAG poisoning vectors and LLM behavioral failure modes | JD lists RAG poisoning as a specific threat vector familiarity preferred qual |
| Summary rewritten to lead with safety evaluation platform credential | JD's first responsibility is translating safety research into guardrails — aeval is the strongest proof point |
| Intuit drift detection bullet reframed as 'proactive regression-detection process analogous to surfacing safety regressions' | JD requires defining evaluation frameworks that surface regressions before they reach customers — this is the closest enterprise analog |
| Intuit language assessment bullet reframed around written communication and executive escalation | JD requires strong written communication translating complex findings for non-technical audiences and knowing when to escalate |
| IBM bullet reframed around escalation judgment and engineering foundations for credible researcher engagement | JD requires technical depth sufficient to engage credibly with safety researchers |
| De Anza Ethical Hacking and Digital Forensics courses highlighted in teaching | Trust and safety / adversarial thinking background is a preferred qual; these courses signal that orientation |
| NeurIPS publication moved to Additional Information as a credential anchor | Establishes research credibility without consuming prime resume real estate now that projects section leads |
JD analysis (20 key phrases)
Key phrases: safety researchmodel behaviorred-teamingevaluation frameworksguardrails and intervention mechanismsagentic AI systemsprompt injectionjailbreaksRAG poisoningmisuse patternssafety product roadmapadversarial inputsbehavioral researchsafety propertiesenterprise AINorth platformsafety reviewthreat vectorsmodel evaluationszero-to-one processes
Hard requirements:
- 5+ years PM or research operations experience
- Technical depth to engage with safety researchers
- Genuine interest in AI safety and model behavior
- Comfortable with ambiguity and unexpected research findings
- Cross-functional alignment across researchers, engineers, product teams
- Strong written communication translating complex model behavior for non-technical audiences
Preferred qualifications:
- Hands-on LLM evaluation, red-teaming, safety benchmarking, or behavioral research
- Familiarity with prompt injection, jailbreaks, RAG poisoning, agentic misuse patterns
- Background in trust and safety, content policy, or research-adjacent operational role
- Experience building zero-to-one processes in research or safety contexts
- Prior exposure to agentic AI systems and unique safety challenges
Per-role mapping (11 roles scored)
| role | score | reframe angle | JD phrases that map |
|---|---|---|---|
| Streamio AI — Founder & CEO | 3/5 | Agentic AI system builder with firsthand exposure to multi-agent orchestration safety challenges and LLM behavioral surface area | agentic AI systems, zero-to-one processes, model behavior, guardrails |
| Fintellect AI — Founder & CEO | 2/5 | RAG pipeline builder with structured output validation — relevant to RAG poisoning and misuse pattern awareness | RAG, agentic AI systems, enterprise AI |
| Intuit — Staff PM | 4/5 | Enterprise platform PM who built proactive regression detection and cross-functional alignment processes at massive scale | evaluation frameworks, safety review, threat vectors, enterprise AI, guardrails |
| Splunk — Senior PM | 3/5 | Technical PM with structured prioritization frameworks and enterprise customer safety/compliance experience | evaluation frameworks, enterprise AI, safety product roadmap |
| Kaiser Permanente — SOA Technical PM | 2/5 | Regulated-domain PM with operational rigor in high-stakes data environments | enterprise AI, safety review |
| IBM — Software Engineer | 1/5 | Engineering foundation supporting technical credibility | — |
| Bank of America — Tech MBA Associate | 1/5 | Quantitative risk modeling background | — |
| RL Workbench | 4/5 | Hands-on RL evaluation platform builder with reward function benchmarking — maps directly to safety evaluation and RLHF alignment research | evaluation frameworks, model behavior, behavioral research, red-teaming |
| aeval — AI Model Evaluation Platform | 5/5 | Built production AI safety evaluation platform with adversarial testing, refusal detection, and automated safety gates — directly maps to Cohere's safety evaluation mandate | evaluation frameworks, red-teaming, safety benchmarking, adversarial inputs, safety properties, model behavior, guardrails and intervention mechanisms, safety review |
| AutoEval | 2/5 | Automated evaluation pipeline with structured safety-style reporting | evaluation frameworks, adversarial inputs |
| BRAIN — Protein Structure Prediction | 2/5 | NeurIPS-published ML researcher with deep neural architecture foundations | behavioral research, model behavior |
Tailored summary
Technical PM and NeurIPS-published researcher with 12+ years bridging AI research and product delivery — uniquely positioned to translate safety research findings into guardrails and product-level controls. Built aeval, a production AI safety evaluation platform with adversarial testing, refusal detection, automated safety gates, and statistical regression detection. Hands-on experience with agentic AI systems, RAG pipelines, multi-agent orchestration, and RL post-training benchmarking across TRL, VeRL, OpenRLHF, and NeMo RL. Scaled enterprise AI platforms to 675M+ engagements at Intuit, building cross-functional alignment processes and proactive drift/regression detection programs.