← baseten / Product Manager, Developer Experience
brief / art_DlGvjVvjaPg
role
model
anthropic/claude-sonnet-4.6
created
2026-06-11T17:22
Company snapshot
Baseten is an AI inference infrastructure company that enables engineering teams to deploy, serve, and scale ML models in production. Their platform is used by high-velocity AI-native companies including Cursor, Notion, Abridge, Clay, Gamma, and Writer for mission-critical inference workloads. They recently closed a $300M Series E (backers include BOND, IVP, Spark Capital, Greylock, and Conviction), signaling aggressive growth and platform expansion. Baseten has a strong engineering-first culture and is actively building out its product function — PMs are expected to be deeply technical and customer-embedded. Their core product surfaces include a CLI, SDKs, a model deployment console, Truss (their model packaging framework), and Chains for multi-model composition; specific recent internal roadmap details are not publicly available.
Team stack
Based on the JD and public signals: Python-first SDK and CLI tooling (likely Click or Typer for CLI, PyPI-distributed); Truss as the model packaging layer (open-source, Python/YAML config); model serving likely built on top of vLLM, TensorRT-LLM, or custom triton-based engines (inferred from 'choosing a serving engine' language in JD); infrastructure likely on AWS/GCP with Kubernetes orchestration (based on inference-at-scale positioning); console likely React/TypeScript frontend; deployment config via config.yml / model.py / chains.py (explicitly named in JD); CI/CD and GitOps patterns for environment promotion (dev→staging→prod, canary/shadow/A/B mentioned explicitly); MCP/agent-first API design emerging as a first-class concern (JD calls out 'agent-driven ways developers build'). Internal observability and telemetry stack unknown.
Likely questions (10)
| area | question | why |
|---|---|---|
| system_design | Walk us through how you would design the end-to-end developer journey from 'model runs on my laptop' to 'serving production traffic' — what are the key friction points and how would you instrument and reduce them? | The JD explicitly states this as the core product mission: 'go from it runs on my laptop to it's serving production traffic in minutes, on their own.' They want to see systems thinking about onboarding + deployment lifecycle. |
| system_design | How would you design a CLI and SDK that treats an AI coding agent (e.g., Claude Code, Cursor) as a first-class user alongside a human developer — what changes in the API contract, error surfaces, and output formats? | JD calls out 'agent-driven ways developers build' and 'treat the coding agent as a first-class user, not a feature' as explicit requirements. This is a differentiating design philosophy they want to probe. |
| domain | Baseten uses Truss for model packaging. If you were redesigning the config.yml / model.py authoring experience to reduce first-deploy errors, what would you change and why? | Truss and config surfaces are explicitly named in the JD as owned surfaces. They want to know if you've actually deployed models and have opinions on packaging DX. |
| domain | Describe how you'd build a progressive delivery system for ML models — canary, shadow, A/B, and rollback. What are the ML-specific failure modes that make this different from standard software releases? | JD explicitly lists 'canary, shadow, A/B, scheduled promotion, and one-gesture rollback' as owned surfaces. ML model releases have unique failure modes (latency regression, output quality drift) that differ from code deploys. |
| behavioral | Tell me about a developer-facing product you shipped — SDK, CLI, or API — where you personally wrote code. What did you build, what tradeoffs did you make, and what would you do differently? | JD requires 'built dev tools as an engineer before you became a PM' and 'shipped SDKs, CLIs, APIs with your own hands.' This is a hard filter question. |
| behavioral | Describe a time you turned direct developer/customer feedback into a roadmap decision that surprised your engineering team or changed the direction of a platform product. | JD emphasizes 'Voice of the Developer — you live in customer conversations' and 'lets customer reality drive the call.' They want evidence of customer-obsessed PM behavior, not just strategy. |
| coding | You're reviewing a PR for a new baseten deploy CLI command. The engineer has returned a generic 'DeploymentError: failed' on build failure. Walk us through how you'd spec the error handling and output format for both human and agent consumers. | JD calls out 'errors surface and builds show progress' as a first-run experience requirement. Error design for dual human/agent consumers is a concrete technical taste signal. |
| culture | Baseten's product function is nascent — you'd be one of the first PMs. How do you earn trust and influence roadmap in a company where engineers have historically owned product direction? | JD explicitly states 'PMs at Baseten don't sit above engineers — you earn ownership by being technical.' This is a culture-fit and self-awareness question about operating in an eng-led org. |
| domain | You're building a Model Library / model discovery experience. What signals would you use to surface the right model to a developer who just signed up, and how does that change for an agent making the same request programmatically? | JD lists 'model discovery and deploy via the Model Library' as an owned surface and agent-first design as a core philosophy. Tests both UX thinking and API-first design instincts. |
| behavioral | At Intuit you scaled ICE to 675M engagements and reduced onboarding from weeks to minutes. What was the hardest technical constraint you had to understand deeply to make that happen, and how did you develop that understanding? | Baseten will probe the Intuit platform story as the closest analog to their own scale and developer-onboarding mission. They want to know if the candidate understands the infrastructure, not just the metrics. |
Talking points
- At Intuit, I owned the ICE Self-Service platform end-to-end — DevPortal, GitOps config, and the ICE Playground — and cut developer onboarding from 2–3 weeks to minutes in pre-prod and under 24 hours for production. I also extended Java and Python SDK Starter Kits with scaffolding, Gradle/Maven build configs, and CI/CD integration so developers could go from zero to production-ready microservice in minutes. That's the same 'laptop to production' problem Baseten is solving, and I've done it at 675M-engagement scale.
- I've shipped developer tooling with my own hands, not just managed it. I built the RL Workbench — a full post-training platform covering GRPO/DPO across TRL, VeRL, OpenRLHF, and NeMo RL with live SSE metric streaming, GPU Docker passthrough, and 12 RL algorithm implementations. I also built aeval, a local-first model evaluation platform with a FastAPI orchestrator, TimescaleDB, Redis job queue, and a Next.js dashboard. I understand what it means to deploy and serve models because I've done it, including on Apple Silicon MPS and CUDA.
- I've designed and shipped multi-agent orchestration architecture: the OpenClaw framework in StreamIO implements a gateway protocol, subagent delegation, profile management, and session switching across multiple industry verticals. This maps directly to Baseten's Chains surface and the JD's call for evolving DevEx for developers building multi-modal agents — I have concrete opinions on agent-first API contract design from having built one.
- At Splunk, I owned Search Service (Go microservices), Search Catalog (PostgreSQL metadata), and SPL/SPL2 — and delivered the Scheduler Service end-to-end in ~4 months. I've lived in the infrastructure layer, written PRDs for microservice backlogs, and achieved up to 10x query performance improvements for beta customers. I can earn engineers' trust because I've been under the hood of distributed systems, not just above them.
- I have a NeurIPS-published paper on neural networks for protein structure prediction (2014), and I recently rebuilt that system in PyTorch spanning 413 parameters to 8B — a 19-million-fold scale increase — with MLflow, Optuna HPO, FastAPI serving, and 823 automated tests. My ML depth is not surface-level; I can reason about model serving tradeoffs, hardware sizing, and inference optimization with Baseten's engineering team as a peer.