← brex / Senior Product Manager, AI
cover_letter / art_UYuoxz68IKE
Cover letter
Dear Brex Hiring Team,
Brex is building the financial OS that lets companies move faster — not by adding another dashboard, but by eliminating the manual work that slows finance teams down. That mission resonates directly with work I've done: at Fintellect AI, I built a mobile-first AI financial platform from prototype to App Store, architecting RAG pipelines, multi-provider LLM orchestration, and domain-specific AI advisors that delivered context-aware guidance at the transaction level. The problem Brex is solving for finance admins — continuous compliance and real-time spend intelligence — is the same problem I've been building toward from the fintech side.
**Technical and AI Foundation**
My AI work spans from the foundational to the applied. In 2004, I hand-coded backpropagation through time in C++ for a protein structure prediction system that was accepted at NeurIPS 2014. In 2026, I rebuilt that system in PyTorch across five architectures — feedforward, GRU, Transformer, ESM-2, and multi-task — scaling from 413 to 8 billion parameters with MLflow tracking, Optuna hyperparameter optimization, and 823 automated tests. That arc matters because it reflects how I engage with AI systems: not as a consumer of tooling, but as someone who understands what's happening underneath.
More directly relevant to this role: I built aeval, a local-first model evaluation platform with five core eval types (factuality, reasoning, instruction-following, safety, code generation), adversarial safety testing with refusal detection, bootstrap confidence intervals, Welch's t-test, and Cohen's d effect size. I also built an RL post-training workbench covering the full RLHF/DPO pipeline — Reward Lab for A/B testing reward functions across GSM8K, MATH, HumanEval, and UltraFeedback; a Playground for live GRPO/DPO training with SSE metric streaming; and an Arena for head-to-head benchmarking across TRL, VeRL, OpenRLHF, and NeMo RL. I've implemented 12 RL algorithms with standardized throughput, memory, and convergence benchmarking. This is the kind of eval infrastructure and iterative quality work the JD describes — designing evals that prove an agent works, then rebuilding them as the models shift underneath.
On the agentic side, I built OpenClaw, a multi-agent orchestration framework with a gateway protocol, subagent delegation, profile management, and session switching — coordinating AI workflows across real estate, insurance, health/dental, and financial-markets verticals. I also shipped Vantage, an AI job search platform with 40+ agent tools, enforced mutation approvals, streaming agent chat over SSE, and an automated AI judge for behavioral interview rehearsal. These aren't demos — they're production systems with real users, real failure modes, and real iteration cycles.
**Why This Role**
The Finance Admin AI work at Brex is exactly the kind of problem I want to own: an agent operating in a trust-critical context, where "done" is a quality bar that keeps moving, and where the PM's job is to define what good judgment looks like before the engineering team can build it. The auditor and business analyst framing in the JD maps precisely to the two hardest problems in agentic product design — when should the agent act versus recommend, and how do you evaluate an answer that has no single correct value. I've lived both of those questions building Fintellect and Vantage, and I want to go deeper on them at Brex's scale.
**Role-Specific Connection**
The compliance auditor agent — reading a company's own expense policy, reviewing every transaction, gathering missing context from employees, grouping exceptions, and surfacing patterns — is a system I understand structurally. At Fintellect, I built 13 specialized AI advisors with structured-output validation, fallback routing, and context-aware advisory on live Alpaca market data. The challenge of making an agent's policy judgment trustworthy enough to act on, rather than just flag, is one I've approached through eval design and incremental trust-building — exactly the framework the JD describes. On the analyst side, I've shipped natural-language query interfaces over financial data with real-time market insights and AI-extracted sentiment, and I understand what it takes to model spend data so the answers are actually right.
**Selected Prior Experience**
- **Fintellect AI:** Architected a RAG retrieval pipeline (ChromaDB) with multi-provider LLM orchestration (Claude, GPT-4, Gemini), fallback routing, structured-output validation, and token-budget optimization; built 13 domain-specific AI advisors delivering guided, context-aware financial advisory on live market data.
- **aeval:** Built a model evaluation platform with adversarial safety testing, refusal detection, statistical rigor (bootstrap CIs, Welch's t-test, Cohen's d), CI/CD regression detection, and automated safety gates — directly applicable to designing evals for agents operating in financial contexts.
- **OpenClaw / Vantage:** Built multi-agent orchestration framework with gateway protocol and subagent delegation; shipped production agentic platform with 40+ tools, enforced mutation approvals, and an automated AI judge — with real iteration on where agents should act versus stop and ask.
- **RL Workbench:** Implemented 12 RL algorithms with standardized benchmarking across TRL, VeRL, OpenRLHF, and NeMo RL; built eval infrastructure for reward function A/B testing across GSM8K, MATH, HumanEval, and UltraFeedback — experience directly applicable to iterating agent quality as underlying models shift.
- **Intuit — ICE Platform:** Scaled platform to 675M+ engagements in FY23 across QuickBooks, TurboTax, Mint, Mailchimp, and Credit Karma; drove 275% YoY engagement growth; scaled throughput from 6K to 50K TPS via rSocket migration supporting ~1.5M concurrent connections with sub-25ms TP99 — evidence of platform product ownership at enterprise scale.
- **Intuit — Developer Tooling:** Conducted enterprise-wide Service Language Assessment across 9 languages, analyzing usage data and developer feedback to inform CTO-level strategic decisions; worked closely with telemetry and usage data (SQL, BigQuery) to prioritize developer pain points across ~20 mobile apps and 30+ product SKUs.
- **Splunk — Search Orchestration:** Owned Search Service (Go microservices), Search Catalog (PostgreSQL metadata), and SPL/SPL2; delivered Scheduler Service end-to-end in ~4 months; led query performance optimization achieving up to 10x improvements — experience with probabilistic, optimization-driven systems where the work is making it steadily better.
**Closing**
Brex's bet is that every finance admin deserves an auditor and an analyst they can't hire — agents that run on every transaction, answer questions in plain language, and earn enough trust to act. That's not a feature; it's a new category of financial infrastructure. I've spent the last several years building toward exactly this intersection of agentic AI, financial data, and trust-critical product design. I'd welcome the chance to bring that work to Brex and help define what this category looks like.
Thank you for your consideration.
---
**O. Felix Amoruwa**
famoruwa@berkeley.edu | 909-731-9011 | felixamoruwa.info