← databricks / Staff Product Manager, Agentic AI Applications
brief / art_yD84KspZSyE
role
model
anthropic/claude-sonnet-4.6
created
2026-06-24T23:22
Company snapshot
Databricks is the data and AI company founded by the original creators of Apache Spark, Delta Lake, MLflow, and the Lakehouse architecture, serving 10,000+ organizations including over 50% of the Fortune 500. The company has been on an aggressive expansion trajectory in AI, acquiring MosaicML (2023) to bolster LLM training capabilities and investing heavily in its Data Intelligence Platform. Databricks is building an internal Agentic Enterprise Applications Platform — a governed, multi-tenant agent runtime enabling GTM, Finance, HR, Legal, and Product teams to ship production-grade agentic applications in weeks. Engineering reputation is strong: the company is known for open-source leadership (Spark, Delta, MLflow), rigorous technical culture, and deep integration between data infrastructure and AI/ML tooling. Specific recent internal org details are not publicly available; inferences below are based on the JD.
Team stack
Based on the JD and Databricks' public engineering signals: agent runtime likely built on or alongside LangGraph/LangChain-style orchestration with durable execution semantics (likely Temporal or a custom checkpoint layer); MCP (Model Context Protocol) as the connector standard for enterprise systems of record; Unity Catalog for governance and identity propagation; MLflow for evaluation pipeline instrumentation and experiment tracking; Delta Lake / Lakehouse for the intelligence/context layer; vector search (likely Databricks Vector Search), knowledge graphs, and structured SQL sources for unified retrieval; Python-first SDK/CLI developer experience; CI/CD evaluation gates likely integrated with GitHub Actions or Databricks Workflows; model gateway abstracting OpenAI, Anthropic, and internal models. Frontend component library stack is uncertain. Infrastructure is cloud-agnostic but AWS/Azure/GCP multi-cloud based on Databricks' customer base.
Likely questions (10)
| area | question | why |
|---|---|---|
| system_design | Walk us through how you would design the three-layer intelligence architecture (knowledge graph, context graph, temporal memory) for the Agentic Platform. How do you ensure unified retrieval with source traceability across vector, structured, and graph sources? | The JD explicitly calls out this three-layer architecture as a core platform responsibility. Databricks will probe whether you can translate this from concept to concrete data architecture decisions. |
| system_design | How would you design a managed agent runtime that supports multi-step orchestration with durable execution, model gateway abstraction, and per-agent guardrails (cost ceilings, blast radius limits)? What are the key trade-offs between sync and async execution models here? | The JD lists 'define and drive the agent runtime' as the first impact area and explicitly calls out sync vs. async as a trade-off engineers will discuss with the PM. |
| domain | You've built an MCP SDK integration in StreamIO. How would you design a connector SDK that lets domain teams (Finance, HR, Legal) onboard new systems of record without requiring platform-team involvement? What are the hardest problems in identity propagation and idempotency for enterprise connectors? | MCP connector ecosystem ownership is a named JD responsibility, and the candidate has direct MCP SDK experience to draw from. |
| domain | Describe how you would design an AI-judge evaluation pipeline with offline golden datasets, online LLM-as-judge scoring, domain-specific judges, and mandatory CI/CD gates. What does 'no agent reaches production without passing quality thresholds' actually mean operationally? | The JD explicitly names evaluation as a core ownership area and states it is 'the hardest part.' Databricks will test depth here. |
| behavioral | Tell me about a time you owned a platform roadmap that 4+ other teams depended on. How did you prioritize competing requests from domain teams without direct authority over any of them? | The JD requires 'proven ability to lead cross-functional initiatives across 4+ teams without direct authority' — this is a named requirement and a classic Staff PM signal. |
| behavioral | Describe a developer experience you shipped — SDK, CLI, templates, or self-service workflow. How did you measure success beyond feature delivery? What was your developer NPS or adoption curve story? | The JD explicitly states 'you measure success by adoption and developer NPS, not feature count' and lists SDK/CLI/templates as required experience. |
| coding | Given a multi-agent workflow where Agent A calls Agent B which calls a CRM connector — how would you instrument observability and tracing so a domain team can debug a failed run in production? What data would you capture at each hop? | Agentic runtime observability is implied by the governance and reliability requirements in the JD; this tests whether the candidate can think like an engineer-adjacent PM on deeply technical instrumentation problems. |
| culture | Databricks is building this platform for internal teams first — GTM, Finance, HR — before potentially externalizing it. How do you think about the tension between moving fast for internal customers versus building a platform that can generalize? Where do you draw the line? | The JD describes an internal-first platform with a federation model (platform-built → domain-built → citizen developer). This is a culture/judgment question about platform vs. product thinking. |
| domain | What is your mental model for the difference between a platform and an application in the context of agentic AI? How does that distinction change your API design decisions and your definition of 'done'? | The JD explicitly states 'you understand the difference between a platform and an application, and you have opinions about API design' — this is a direct filter question. |
| behavioral | Tell me about a time you had to write a strategy document or PRD that drove alignment at the VP or CIO level on a technically ambiguous AI/ML initiative. How did you handle disagreement on technical direction? | The JD calls out 'strategy documents, PRDs, and executive briefs that drive alignment at VP and CIO level' as a hard requirement for this Staff-level role. |
Talking points
- Built OpenClaw multi-agent orchestration framework from scratch at StreamIO (ev_9WxED1-HP5E): implemented gateway protocol, subagent delegation, profile management, and session switching — directly analogous to the managed agent runtime and MCP connector ecosystem Databricks is building. Can speak concretely to identity propagation, tool invocation governance, and multi-agent coordination patterns.
- Built aeval, a production AI evaluation platform (ev_HFOZxjussHw): 5 eval types, LLM-as-judge scoring, adversarial safety testing, bootstrap confidence intervals, Welch's t-test, Cohen's d, and CI/CD regression gates with automated safety thresholds — this is the exact evaluation and quality framework the JD describes as a core ownership area, built and shipped independently.
- At Intuit (Staff PM, Developer Frameworks & Platform Infrastructure), delivered ICE Self-Service platform that reduced developer onboarding from 2–3 weeks to minutes, scaled to 675M+ engagements at 50K TPS, and extended Java/Python SDK Starter Kits with scaffolding, CI/CD integration, and testing frameworks — direct evidence of owning a developer platform that other teams build on at enterprise scale.
- Built RL Workbench benchmarking GRPO/DPO across TRL, VeRL, OpenRLHF, and NeMo RL (ev_jPb9hXIN--w), implementing 12 RL algorithms with standardized throughput/memory/convergence benchmarking — demonstrates the technical depth to engage credibly with Databricks engineers on model training, evaluation pipelines, and AI infrastructure trade-offs at a level rare for PMs.
- NeurIPS 2014 published researcher on neural networks for protein structure prediction, with a 2026 rewrite spanning 413 to 8B parameters using PyTorch, MLflow experiment tracking, and Optuna HPO — establishes long-arc AI/ML credibility and direct familiarity with MLflow, a core Databricks open-source asset named in the JD's nice-to-haves.