jobsearch v0.0.1

← togetherai / Forward Deployed Engineer (Inference & Post-Training)

interviewer_questions / art_5by7BVviqjI

role
togetherai / Forward Deployed Engineer (Inference & Post-Training)
model
anthropic/claude-sonnet-4.6
created
2026-10-08T16:42

Interviewer

Rochelle Mattern is Head of Field Engineering at Together AI, joining in September 2025. Her background spans field engineering leadership at SambaNova Systems (a direct competitor in AI inference hardware/software), Forethought (agentic AI), and Google Cloud (enterprise customer engineering). She built and scaled pre-sales and post-sales technical teams, owns the product-field feedback loop, and has a track record of reducing time-to-value through structured onboarding and POC frameworks. As the hiring manager or senior leader for this FDE role, she will probe deeply on customer-facing technical depth, ability to run complex POCs, inference/post-training hands-on credibility, and fit within a field engineering culture that bridges sales, engineering, and product.

My profile through their lens

From Rochelle's perspective, Felix is a rare candidate who combines genuine ML research depth (NeurIPS, hand-coded BPTT, RL workbench) with platform-scale PM experience (Intuit 675M engagements, Splunk microservices) and recent hands-on builder credibility (12-algo RL workbench, aeval, multi-agent orchestration). She will be impressed by the RL post-training breadth — GRPO/DPO/PPO across TRL/VeRL/OpenRLHF/NeMo RL — which maps directly to the role's post-training requirements. However, she will probe hard on inference engine depth (vLLM, TensorRT-LLM, SGLang) since the resume does not explicitly name these tools, and on customer-facing field engineering experience, which is largely absent — Felix's background is founder/PM/researcher, not pre-sales SE or FDE. She will also want to understand whether Felix can operate in a revenue-impacting, customer-aligned role rather than a purely internal platform or product role.

Questions they may ask (21)

categoryquestionwhyhow to prepare
resume_deep_dive Walk me through your RL Workbench — specifically the Arena component where you benchmark TRL, VeRL, OpenRLHF, and NeMo RL. What were the most meaningful performance differences you observed across frameworks, and what would you tell a customer trying to choose between them for a GRPO run on a 70B model? The RL Workbench is Felix's strongest direct evidence for the post-training requirement. Rochelle came from SambaNova where she owned technical wins on fine-tuning of open-source models. She will want to verify this is real, production-level work and that Felix can translate benchmark findings into opinionated customer guidance — a core FDE skill. Prepare a crisp 2-minute narrative on Arena architecture and GPU Docker passthrough setup, then have concrete benchmark findings ready (throughput, memory, convergence deltas across at least 2 frameworks). Practice framing the answer as a customer recommendation, not just a technical description.
resume_deep_dive At Intuit you scaled ICE to 675M engagements and 50K TPS via rSocket migration. What were the actual inference bottlenecks you hit at scale, and how did you diagnose and resolve them? How does that experience translate to optimizing LLM inference throughput today? Rochelle will want to test whether Felix's platform scale experience is genuinely technical or primarily PM-level. The rSocket/1.5M concurrent connections detail is specific enough to probe. This also bridges to the inference optimization requirements in the JD. Prepare to go deep on the rSocket migration: what the bottleneck was (connection overhead, serialization, head-of-line blocking), how it was diagnosed, and what the tradeoffs were. Then explicitly bridge to KV cache, batching, and token throughput in LLM inference contexts.
resume_deep_dive Your aeval platform implements bootstrap confidence intervals, Welch's t-test, and Cohen's d for model evaluation. How would you use aeval — or a similar evaluation framework — to help a Together AI customer validate that a fine-tuned model is actually better than the base model before they move to production? Rochelle's SambaNova role involved proof-of-value planning and execution. She will want to see whether Felix can translate his evaluation tooling into a structured POC/POV framework for customers — a direct analog to her own field engineering methodology. Prepare a concrete POV framework: baseline eval → fine-tune → eval delta with statistical significance → production gate. Be ready to discuss what metrics matter for different customer use cases (latency, accuracy, safety) and how to set pass/fail thresholds.
resume_deep_dive You've built multi-agent orchestration (OpenClaw) and RAG pipelines (Fintellect) as a founder. How do these architectures interact with inference optimization concerns — specifically, how does multi-agent orchestration change your KV cache strategy or batching approach compared to single-turn inference? The JD requires inference engine optimization in the context of production AI teams. Felix's multi-agent work is evidence of real deployment experience, but Rochelle will want to verify he understands the inference implications of agentic workloads — a nuanced area that separates deep practitioners from surface-level builders. Study how multi-turn, multi-agent workloads affect KV cache reuse (prefix caching, radix attention in SGLang), batching strategies (continuous batching vs. static), and latency SLAs. Prepare a concrete example from OpenClaw where inference configuration choices mattered.
technical_domain Together AI's platform supports vLLM, TensorRT-LLM, and SGLang. A strategic customer is running Llama-3.1-70B for a high-throughput batch inference workload (10K requests/hour, latency-insensitive) vs. a real-time agentic use case (sub-200ms TTFT, low concurrency). Walk me through how you'd configure the inference engine differently for each workload. This is the core technical competency for the role. Rochelle came from SambaNova where inference engine selection and configuration was central to field wins. The JD explicitly lists KV cache tuning, speculative decoding, tensor parallelism, and quantization. Felix's resume does not explicitly name vLLM/SGLang/TensorRT-LLM, so this is a direct probe of that gap. Study vLLM's chunked prefill, continuous batching, and prefix caching; SGLang's RadixAttention and speculative decoding; TensorRT-LLM's static batching and INT8/FP8 quantization. Prepare a concrete configuration comparison table for batch vs. real-time workloads covering engine choice, parallelism strategy, KV cache settings, and quantization.
technical_domain A customer wants to fine-tune Mistral-7B using GRPO for a reasoning task. They have 8x H100s and a dataset of 50K prompt-completion pairs with verifiable rewards. Walk me through the system design decisions you'd make: framework selection, LoRA vs. full fine-tune, batch size, gradient accumulation, reward model architecture, and how you'd monitor training stability. This maps directly to Felix's RL Workbench work and the JD's post-training requirements. Rochelle will use this to validate that Felix can give opinionated, production-grade guidance to customers — not just describe what algorithms exist. Prepare a concrete GRPO system design: TRL vs. VeRL for 8xH100, LoRA rank selection for 7B, KL penalty tuning, reward normalization, and key instability signals (reward hacking, entropy collapse). Reference your RL Workbench benchmark findings to ground the recommendations.
technical_domain Explain speculative decoding to me as if I'm a customer's ML engineer who has heard the term but never implemented it. Then tell me when you would and would not recommend it for a Together AI customer. Speculative decoding is explicitly listed in the JD. Rochelle's field engineering background means she values the ability to explain complex concepts clearly to customers at varying technical levels — a core FDE communication skill she will assess. Prepare a clean 90-second explanation of speculative decoding (draft model + verifier, token acceptance rate, latency gain). Then prepare a decision framework: when it helps (low-to-medium concurrency, latency-sensitive, small draft model available) vs. when it doesn't (high concurrency, throughput-bound, no suitable draft model).
technical_domain What's your mental model for choosing between INT4, INT8, FP8, and FP16 quantization for a customer deploying a 70B model on A100s vs. H100s? What are the accuracy tradeoffs you'd flag upfront? Quantization strategy is explicitly listed in the JD under inference optimization. This is a standard technical screen for inference-focused roles and will reveal whether Felix has hands-on experience or only theoretical knowledge. Study GPTQ, AWQ, and GGUF for INT4/INT8; FP8 native support on H100 vs. emulation on A100; perplexity degradation benchmarks for popular models. Prepare a decision matrix covering memory footprint, throughput gain, accuracy loss, and hardware compatibility.
gap_transition Your background is primarily as a founder and PM — you've built products and platforms, but you haven't held a customer-facing field engineering or solutions engineering role. How do you think about the transition from internal platform builder to external customer partner, and what do you think will be hardest about it? This is the most significant gap relative to the role. Rochelle has spent her entire career in field engineering (Google CE, SambaNova SE Director, Forethought SE/PS). She will probe whether Felix has realistic self-awareness about this transition and a credible plan. Acknowledge the gap directly and honestly. Frame your customer discovery work at Streamio/Fintellect as proto-field-engineering. Emphasize that your technical depth (RL workbench, aeval, inference experience) means you can add value immediately on the hardest POCs. Prepare a specific example of translating technical complexity into customer-facing guidance.
gap_transition The JD mentions expert-level hands-on experience with vLLM, TensorRT-LLM, or SGLang. Your resume references inference work but doesn't name these engines explicitly. What's your actual hands-on experience with these specific tools, and how quickly could you get to expert level? Rochelle will have seen many candidates who know the theory but haven't run production inference engines. As someone who ran technical wins at SambaNova (a competing inference platform), she knows exactly what expert-level looks like and will probe this gap directly. Be honest about your current level with each engine. If you've used vLLM via Ollama or local deployments, say so specifically. Prepare a concrete 30-day ramp plan: deploy vLLM locally, run benchmarks, tune KV cache, test speculative decoding. Demonstrating structured learning velocity is more credible than overclaiming.
gap_transition This role has a significant revenue impact component — you're supporting strategic accounts and influencing POC outcomes that directly affect sales. You've been a founder and PM, but not in a quota-carrying or revenue-aligned field role. How do you think about operating in that context? Rochelle is a President's Club winner from Google and has led revenue-generating field teams. She will want to assess whether Felix understands and is motivated by the commercial dimension of FDE work, not just the technical side. Prepare examples from Intuit where your technical decisions had direct revenue impact (ICE Presence $480K/month, $1M+ opex mitigation). Frame your founder experience as inherently revenue-aligned. Show genuine enthusiasm for the customer success dimension, not just the engineering work.
behavioral_situational Tell me about a time you had to give a customer or stakeholder a technically honest answer that they didn't want to hear — for example, that their proposed architecture wouldn't scale, or that a feature they wanted wasn't the right solution. How did you handle it? Rochelle's field engineering philosophy at Forethought and SambaNova emphasized opinionated onboarding and honest technical guidance. The JD uses the word 'opinionated' explicitly. She will want evidence that Felix can be diplomatically direct with customers. Prepare a specific STAR story — ideally from Intuit or Splunk where you pushed back on a customer or internal stakeholder on a technical decision. Emphasize the outcome and the relationship preservation, not just the technical correctness.
behavioral_situational Describe a situation where you had to rapidly learn a new technical domain to solve a customer or stakeholder problem. What was your approach, and how long did it take you to get to a credible level? FDE roles require continuous learning as the model landscape evolves rapidly. Rochelle will want evidence of Felix's learning velocity and methodology, especially given the inference engine gap. Use the RL Workbench as a primary example — you went from conceptual knowledge of GRPO/DPO to implementing 12 algorithms across 4 frameworks. Quantify the timeline and describe your learning methodology (papers → implementation → benchmarking).
behavioral_situational Tell me about a time you identified a product gap or platform limitation while working with users or customers, and successfully drove it back into the engineering roadmap. What was your process and what was the outcome? The JD explicitly calls out 'product feedback loop' as a core responsibility. Rochelle built this feedback loop at SambaNova and Forethought. She will want evidence Felix can operate in this bidirectional role between customers and product/engineering. Use the Intuit ICE Self-Service or MSaaS Drift Detection examples — you identified developer pain points via telemetry, built the case, and shipped the solution. Frame it as a customer-to-roadmap loop, not just internal PM work.
behavioral_situational Give me an example of a time you managed multiple high-priority technical workstreams simultaneously — ideally with competing deadlines and different stakeholders. How did you prioritize and what did you deprioritize? FDEs at Together AI support multiple strategic accounts simultaneously. Rochelle managed global pre-sales teams and will probe for evidence of structured prioritization under pressure. Use your founder experience running Streamio AI and Fintellect simultaneously, or the Intuit period managing ICE, MSaaS, and the language assessment concurrently. Be specific about your prioritization framework and what you explicitly chose not to do.
role_specific_scenario A strategic customer — say, a company like Cursor or Decagon — comes to you with a POC requirement: they need sub-100ms TTFT for a code completion use case at 500 concurrent users on Together's platform. Walk me through how you'd scope, design, and execute this POC from day one. Rochelle built POV frameworks at SambaNova and reduced time-to-value by 75% at Forethought through structured POC processes. She will use this to assess whether Felix can run a structured, outcome-oriented POC — the core FDE motion. Prepare a POC framework: requirements gathering (model, hardware, SLA), baseline measurement, configuration hypothesis (speculative decoding, tensor parallelism, KV cache), A/B test design, success criteria, and escalation path. Reference specific Together AI infrastructure where possible.
role_specific_scenario A customer's DPO fine-tuning run is diverging — their reward model loss is spiking after epoch 2 and the model is producing repetitive outputs. They're on a deadline for a board demo. What's your diagnostic process and what are the first three things you check? This is a realistic FDE scenario that tests both post-training depth and customer pressure management. Felix's RL Workbench gives him direct experience with training instability, and Rochelle will want to see structured diagnostic thinking under pressure. Prepare a DPO instability diagnostic checklist: beta coefficient (KL penalty), reference model drift, dataset quality (length bias, reward hacking), learning rate schedule, and gradient norms. Practice delivering this as a calm, structured response to a stressed customer.
motivation_fit Together AI competes with platforms like SambaNova, Fireworks, Anyscale, and others. Why Together AI specifically, and what do you think Together's differentiated position is in the inference and post-training market? Rochelle came directly from SambaNova — a direct competitor. She will probe whether Felix has done genuine competitive analysis and has a credible, specific answer about Together's differentiation, not a generic 'I love AI' response. Research Together AI's specific differentiators: open model marketplace, fine-tuning + inference in one platform, FlashAttention contributions, pricing model, and customer base (Cursor, Decagon, ElevenLabs). Prepare a crisp comparison to SambaNova (hardware-centric, enterprise) vs. Together (software-first, developer-centric).
motivation_fit You've been a founder for the past year building your own AI products. Why are you choosing to join a company now rather than continuing to build, and why this specific role rather than a PM or engineering role? Rochelle will want to understand the genuine motivation behind this transition. FDE is a specific career path that requires customer-facing orientation and comfort with being a technical partner rather than the decision-maker. She needs to assess whether Felix is choosing FDE or defaulting to it. Prepare an honest, specific answer: what you've learned from founding that makes you want to go deep on inference/post-training as a practitioner, why the customer-facing dimension is genuinely appealing (not just acceptable), and why FDE at Together AI is the right next step rather than a PM role.
unique_to_this_interviewer You came from SambaNova, which is a full-stack AI platform with proprietary hardware. Together AI is software-first and hardware-agnostic. From your experience, what are the biggest technical challenges customers face when moving from a hardware-optimized inference environment to a cloud-based, hardware-agnostic platform like Together's? This question is unique to Rochelle's SambaNova background. It tests whether Felix has thought about the competitive landscape from the customer's perspective and can engage Rochelle as a peer on a topic she knows deeply. It also demonstrates genuine curiosity about her experience. Research SambaNova's RDU architecture and its inference optimization approach. Prepare thoughts on the tradeoffs: hardware-specific optimization vs. portability, vendor lock-in vs. flexibility, and how Together's software-layer optimizations (vLLM, SGLang) compare to hardware-level acceleration. Ask Rochelle for her perspective after sharing yours.
unique_to_this_interviewer You've built and scaled field engineering teams from scratch at multiple companies. What does 'great' look like for an FDE in the first 90 days at Together AI, and what are the most common failure modes you've seen in this role? Rochelle has directly hired and built FDE/SE teams at SambaNova and Forethought. Asking this question signals self-awareness and genuine interest in succeeding in the role on her terms. It also gives Felix intelligence about her evaluation criteria. This is a question to ask Rochelle, not to answer yourself. Prepare it as a genuine question after demonstrating technical depth. Listen carefully — her answer will reveal her actual evaluation criteria and what she's seen fail in this role before.

Preparation priorities

  1. 1. INFERENCE ENGINE DEPTH (highest priority): The resume does not explicitly name vLLM, SGLang, or TensorRT-LLM. Deploy vLLM and SGLang locally this week, run benchmarks, tune KV cache and test speculative decoding. Prepare a configuration comparison for batch vs. real-time workloads. This is the single biggest gap relative to the JD requirements.
  2. 2. POC/POV FRAMEWORK: Rochelle's entire career is built on structured technical wins and time-to-value reduction. Prepare a concrete POC methodology — requirements → baseline → hypothesis → A/B test → success criteria — and practice delivering it as a customer-facing narrative, not an internal PM process.
  3. 3. RL WORKBENCH AS ANCHOR STORY: This is Felix's strongest differentiator. Prepare a crisp, customer-facing narrative that translates benchmark findings (TRL vs. VeRL vs. OpenRLHF) into opinionated recommendations. Practice explaining GRPO training instability diagnostics under time pressure.
  4. 4. CUSTOMER-FACING TRANSITION NARRATIVE: The gap from founder/PM to FDE is real and Rochelle will probe it. Prepare an honest, specific story about why FDE is the right role, what customer-facing experience you do have (customer discovery, Intuit developer community, Splunk .conf demos), and a concrete 30/60/90 day ramp plan.
  5. 5. TOGETHER AI COMPETITIVE POSITIONING: Rochelle came from SambaNova (a direct competitor). Prepare a specific, informed view on Together AI's differentiation vs. SambaNova, Fireworks, and Anyscale. Demonstrate you've used the platform and have genuine opinions about its strengths and gaps.

⚠ Watch-outs