← coreweave / Senior Product Manager, Security & Infra
brief / art_JKnaVqiUW98
role
model
anthropic/claude-sonnet-4.6
created
2026-05-20T22:37
Company snapshot
CoreWeave is a GPU-specialized cloud provider positioning itself as 'The Essential Cloud for AI,' offering high-performance compute infrastructure primarily to AI labs, model trainers, and enterprise AI teams. Founded in 2017 as a crypto-mining operation, it pivoted aggressively to GPU cloud and became publicly traded on Nasdaq (CRWV) in March 2025, one of the largest tech IPOs of that year. The company has grown rapidly on the back of demand from frontier AI labs and is known for its bare-metal GPU clusters, low-latency networking (InfiniBand), and Kubernetes-native orchestration. Engineering reputation is strong among ML infrastructure practitioners — CoreWeave is seen as a serious alternative to hyperscalers for training workloads. Recent moves include significant data center expansion and enterprise customer growth; specific named deals or internal projects are not confirmed here.
Team stack
Based on the JD, the team's core platforms include: Okta (Identity Engine + Identity Governance — explicitly named), Opal (access governance — explicitly named), Google Workspace, Slack, Zoom (SaaS collaboration — explicitly named), MDM/endpoint management tools (likely Jamf for Mac, Intune for Windows — inferred from Mac/Linux/Windows mention), VDI solutions (vendor not specified in JD — likely Citrix, VMware Horizon, or AWS WorkSpaces based on industry norms). Automation tooling likely includes Python and/or Go scripting against vendor APIs, with possible Terraform/Ansible for IaC (listed as preferred). CI/CD and Git are explicitly required. Security compliance frameworks referenced: SOX and SOC 2. The broader IT Infra and Ops org likely runs on a cloud-native, Linux-heavy substrate consistent with CoreWeave's GPU cloud environment.
Likely questions (10)
| area | question | why |
|---|---|---|
| domain | Walk me through how you would design an end-to-end joiner/mover/leaver (JML) lifecycle in Okta, including automated provisioning, RBAC assignment, and de-provisioning triggers. What are the failure modes you'd instrument first? | The JD explicitly calls out 'joiner, mover, and leaver flows' as a primary ownership area and asks for secure, automated, reliable access changes — this is the core domain deliverable of the role. |
| system_design | CoreWeave is onboarding new enterprise SaaS applications rapidly. Design a scalable SSO onboarding framework in Okta — covering intake, integration standards, testing, and deprecation of legacy auth patterns. How do you enforce consistency at scale? | The JD explicitly asks the PM to 'establish standards for SSO integrations and app onboarding and offboarding in Okta' — this tests both technical depth and scalable process thinking. |
| system_design | How would you define SLIs and SLOs for a VDI platform serving engineers who need secure remote access to GPU workloads? What does an incident playbook look like for a VDI outage affecting 500 engineers? | The JD explicitly asks the candidate to 'own the VDI and secure remote access experience as a product surface' and define 'SLIs, SLOs, and incident playbooks.' |
| domain | Describe your experience with Just-in-Time (JIT) and time-bound access patterns. How have you balanced security posture with developer/engineer friction when rolling out privileged access controls? | The JD calls out 'Just-in-Time and time-bound access patterns' explicitly as a design and rollout responsibility, and CoreWeave's AI-lab customer base means engineers likely have elevated access needs. |
| behavioral | Tell me about a time you had to translate a complex security or compliance requirement (e.g., SOX ITGC, SOC 2 control) into a product roadmap item that engineering could actually ship. How did you manage tradeoffs between compliance deadlines and engineering capacity? | The JD requires 'partnering with Security and Compliance on access controls, ITGCs, and IT application controls for frameworks such as SOX or SOC 2, including control design and evidence collection.' |
| behavioral | Describe a situation where you converted a set of ad-hoc IT systems engineer responsibilities into a scalable, measurable product capability. What did you have to change about how the team worked, and how did you measure success? | The JD explicitly states the goal is to 'convert existing Staff IT Systems Engineer responsibilities into scalable, measurable product capabilities' — this is a direct behavioral probe of that exact challenge. |
| coding | You need to audit all Okta users who have not completed a Quarterly Access Review (QAR) and auto-suspend accounts older than 90 days with no manager approval. Walk me through how you'd script this using the Okta API in Python or Go — what edge cases do you handle? | The JD requires 'comfort working with at least one programming or scripting language such as Python or Go' and explicitly references 'QAR completion' as a tracked metric and automation of access governance tasks. |
| domain | How would you approach SaaS license governance across 20+ applications (Google Workspace, Slack, Zoom, etc.) at a fast-growing company? What data would you track, and how would you build a reclamation workflow without alienating end users? | The JD calls out 'asset and SaaS lifecycle management, including inventory visibility, licensing utilization, and access governance for endpoints and applications' as an explicit ownership area. |
| culture | CoreWeave moves fast and has a high tolerance for ambiguity — this role is partly about building product structure where little exists today. How do you operate in an environment where you're simultaneously defining the process and delivering against it? | The JD describes converting engineer-owned responsibilities into product capabilities at a hyper-growth public company; CoreWeave's stated values ('Act Like an Owner,' 'Be Curious at Your Core') and the JD's 'high-growth, cloud-native' preferred experience signal this is a key culture fit probe. |
| behavioral | Tell me about a platform or developer-facing product where you defined and tracked operational metrics (e.g., time-to-access, ticket reduction, coverage rates) and used those metrics to reprioritize the roadmap. What did you learn? | The JD explicitly lists metrics like 'time-to-access for new hires, reduction in access-related tickets, SSO coverage, QAR completion, and control failures' as PM-owned KPIs — this probes whether the candidate has actually done metric-driven product management. |
Talking points
- At Intuit, I owned the ICE Self-Service platform end-to-end — reducing developer onboarding from 2–3 weeks to under 24 hours for production, scaling throughput from 6K to 50K TPS, and reaching 675M+ engagements in FY23. That's exactly the kind of 'convert engineer toil into scalable product capability' motion this role is asking for — I've done it at enterprise scale across 30+ product SKUs.
- I built the ICE Drift Detection and Resolution program at Intuit: wrote a Java JAR library to scan Git repos for configuration drift, partnered with Design on a DevPortal remediation UI, and built the remediation roadmap using OpenRewrite. This maps directly to the JD's ask for automation workflows that remove manual IT and security tasks — I've shipped that kind of infra-as-product work before.
- I built the aeval evaluation platform with a FastAPI orchestrator, Redis job queue, TimescaleDB, and CI/CD integration with automated safety gates — and my RL Workbench benchmarks 12 algorithms across TRL, VeRL, OpenRLHF, and NeMo RL with Docker GPU passthrough. This demonstrates the scripting and systems fluency the JD requires: I can read and reason about Python/Go implementations, understand API tradeoffs, and work credibly alongside engineers on configuration and release decisions.
- At Splunk, I owned three microservice backlogs (Search Service in Go, Search Catalog in PostgreSQL, SPL/SPL2) and delivered the Scheduler Service end-to-end in ~4 months — demoed at .conf19. I've operated in security-sensitive, compliance-adjacent environments and built RICE-based prioritization frameworks that balanced Fortune 500 customer requirements with internal engineering capacity. That rigor translates directly to managing IAM roadmap tradeoffs against SOX/SOC 2 deadlines.
- My NeurIPS-published research and 20+ years of hands-on engineering — from hand-coded BPTT in C++ in 2004 to multi-agent orchestration frameworks today — means I can engage with CoreWeave's engineering team as a genuine technical peer, not just a requirements translator. The JD asks for a PM who can 'understand implementation options and tradeoffs' in Python or Go; I can do that and go deeper when the problem demands it.