Design: Cleveland Kitchen Harness v2 — MCP + Skills (no frontend)
Generated by /office-hours on 2026-05-29 Branch: claude/vigorous-mendeleev-1a67b1 Repo: ERP-Unlocked/ordermatic Status: APPROVED Mode: Startup Builds on: dboone31-erp-unlocked/mini-boone-claude-epic-roentgen-295c38-design-20260517-075959.md (Harness v1 wedge, APPROVED, /revenue-leaks SHIPPED as ENG-195 on 2026-05-28)
Problem Statement
Section titled “Problem Statement”After Revenue Leaks shipped to Cleveland Kitchen, Sam B (ops/forecasting) named three additional pain points to David in a hallway talk, with Mac (owner) actively pushing for the most visible one. The four buckets:
- Outbound order confirmation emails — automated “we received your order” on ingest, “order’s been processed” on WhereFour push. “Biggest thing Mac has been on my ass about.”
- Weekly forecast/allocation workflow — Sam spends his entire week in Excel reconciling 10K cases of pickle red onions vs 20K in orders, then sending the sheet to Min for WhereFour import. Mixes data assembly (computable) with political judgment (“who’s bitching the most” — Sam’s quote).
- Short-supply customer email notifications — already built and working by Min, externally to Ordermatic. Not asked to be replaced.
- Revenue leak detection — already shipped (ENG-195).
The strategic question this session sharpened: how should Ordermatic add new capability to Cleveland Kitchen now that Revenue Leaks is live? The harness v1 design assumed additional Astro panels. David’s decision in this session: no new UI surface. Authenticated MCP server + Temporal-driven product workflows (the customer-facing order emails). Cleveland Kitchen already uses Claude for ops work, primarily in Excel — that’s the consumption surface. Sam stays in his existing allocation sheet; Claude-in-Excel queries the Ordermatic MCP server to fetch live inventory, fill rates, and recent shorts; the evidence packet renders in his sheet; Sam decides. The “blow their socks off” UX is Sam never leaving Excel.
Customer-facing order confirmation emails (“we received your order” / “your order has been confirmed in our system”) are product, not just MCP visibility. They fire automatically from server-side Temporal workflows on extracted_order.received and wherefour.order_pushed.succeeded events. Mac sees real emails going to CK’s customers without anyone opening Claude.
Demand Evidence
Section titled “Demand Evidence”Three-stakeholder expansion signal at customer #2:
- Mac (owner, buyer): pressuring Sam for the customer-facing order email. This is the renewal-buying ask.
- Sam B (ops): “I am literally spending my entire week going through our forecast for next week” — quantified, behavioral, recurring weekly. Already adjusting orders in Excel and emailing Min to import to WhereFour.
- Min (data/ops): asked for and got Revenue Leaks. Built short-supply emails himself, so we know what he reaches for when Ordermatic doesn’t have it yet — and we know what NOT to break.
Three named users at the same account with three distinct asks = the harness thesis (“the thing that fights back, customer-by-customer”) is getting concrete validation at the canonical-shape-of-customer for this product.
Quote that frames the architecture:
“in terms of allocating product to a customer and automating that, it is still manual because you kind of have to go look at one who’s bitching the most.”
Sam is telling us the decision is political-relational, not algorithmic. That makes the durable product not auto-allocation but assembling the full evidence packet so a human decides in 30 minutes instead of 40 hours. Every future capability (allocation, AR aging, the rest of Q2C) is this same shape: surface multi-system facts, let the human apply judgment, log the decision so next time we can suggest. Evidence packet, not judgment engine.
Status Quo
Section titled “Status Quo”| Pain | Current workflow | Pain magnitude |
|---|---|---|
| Order confirmation | None. Customer guesses if their PO landed. Mac fields the calls. | ”Mac on my ass” — soft-quantified, customer-experience-level |
| Allocation | Sam exports forecast → Excel → manual fill-rate/short-history lookups → sends edited sheet to Min → Min imports to WhereFour | ~40 hrs/week of Sam’s time |
| Short-supply emails | Min built it himself, generates batch email “short on X Y Z, impacts PO A B C, weld date Z” | Working, ~0 friction |
| Revenue leaks | Shipped as /revenue-leaks Astro panel last week | (out of scope) |
Target User & Narrowest Wedge
Section titled “Target User & Narrowest Wedge”Primary users this batch:
- Mac sees: customer emails go out automatically. He never logs in. He learns Ordermatic is shipping because customers thank him.
- Sam sees: he opens his existing allocation Excel sheet. From Claude-in-Excel he asks “who should I prioritize for next week’s pickle red onions?”. Claude calls Ordermatic MCP tools (inventory snapshot, open orders, fill rates, recent shorts), renders the evidence packet in cells, Sam decides. AI logs the decision back via
record_allocation_decision. Sam never leaves Excel. That is the unlock. - Min sees: he keeps his existing short-supply email pipeline (per P5). New tools (
get_revenue_leaks,list_failed_emails, audit log) are available to him via MCP from whichever Claude surface he prefers; he doesn’t have to use them but they’re discoverable.
Deferred / not in this batch:
- Min’s short-supply email (P5 — don’t touch what’s working)
- Any auto-allocation logic (P4 — evidence packet, not judgment engine)
- Friction-test of “does Sam actually use an AI client today” — see Open Questions
Narrowest wedge (D2 reframed by David): the forwardable artifact is the customer email itself (Mac’s win), not a dashboard screenshot. Mac shows off Ordermatic by forwarding the customer email he just got from his own customer.
Constraints
Section titled “Constraints”- David is part-time on Ordermatic until $500K raise closes — engineering velocity is the binding constraint.
- Fabio (data eng) has been verbally committed to replicating Sam’s allocation sheet schema into the warehouse. Sheet share is pending.
- Customer-facing copy is internal decision (per
feedback_customer_copy_is_internal.md) — don’t loop CK in for pre-ship email template review; measure reaction post-ship from behavior. - No new frontend surface this batch (D3 decision). Revenue Leaks panel stays; everything new ships as MCP tools + skill definitions + server-side workflows.
- PR base is
staging; tests via Docker; worktree needspnpm install --prefer-offline --ignore-scriptsbefore first commit (per project memory). - Symphony orchestrator is the multi-agent execution substrate David referenced as “my factory” on the call — relevant for build delivery, not for runtime architecture.
Premises
Section titled “Premises”P1 (revised twice): All four CK pain buckets belong inside the same Ordermatic capability surface for CK — but that surface is now an authenticated MCP server + Temporal product workflows, consumed via Claude-in-Excel (Sam’s allocation cockpit) and via server-triggered customer emails (Mac’s renewal artifact). NOT an Astro control tower. Same harness thesis (“the thing that fights back, per customer”), different surface. Revenue Leaks Astro page stays as-is; everything new is MCP-native or workflow-native. Skill files dropped per spec-review scope cut.
P2 (revised after subagent reframe): Order confirmation emails ship before allocation tools, because the customer email itself is the forwardable artifact CK shows off externally. Allocation tools are the churn-prevention artifact (saves Sam ~40 hrs/week); they ship second. Different reasons than the original “buyer drives renewal” framing, same sequencing.
P3: “Order confirmation email” is a 1-week part-time build. Risk: event detection (what counts as “received” vs “processed”), per-customer email mapping when SPS Commerce orders don’t carry direct end-customer email addresses, deliverability/SPF for CK’s domain, suppression toggles. If P3 is wrong by more than 1 week, allocation slides into a third sprint.
P4: Sam’s allocation Excel encodes the data-assembly side of the decision (fill rates, recent shorts, available qty). It does NOT encode “who’s bitching the most” — that’s Sam’s tacit judgment. MVP MUST be data-assembly only. Auto-allocation in v1 = scope blowup + erodes trust with Sam.
P5: Min’s working short-supply email pipeline does NOT get touched in this batch. He didn’t ask for it to be replaced; touching it is all downside (regression risk, no upside). Centralize later, after we earn another expansion.
P6: Three named CK stakeholders with three distinct asks = account-expansion signal. CK is now the canonical shape of customer for this product. The MCP+skills surface should be designed such that the second CK-shaped customer (Genfit on WhereFour, or next signed WhereFour shop) gets the same tools by configuration, not by code.
P7 (new — load-bearing for the MCP architecture): Customer emails MUST trigger from server-side workflows (Temporal), not from Sam-querying-his-AI. Auto-execution belongs on the server; MCP tools exist for visibility, manual override, and template-edit. Without P7, no email ever sends unless an operator opens Claude.
P8 (new — resolves prior OQ1): Cleveland Kitchen already uses Claude for ops work, with Excel as the primary surface. The MCP server’s primary consumer is Claude-in-Excel, NOT Claude Desktop. Desktop is a fallback for workflows that don’t fit in a sheet (revenue-leak triage, support investigations). This is a confirmed user behavior, not an architectural hope.
Evidence test that P1+P8 (Claude-in-Excel + MCP) is wrong: within 14 days of MCP server going live to CK, if Sam has not called Ordermatic MCP tools from his allocation sheet AT LEAST ONCE on a real allocation cycle, then either (a) Claude-in-Excel doesn’t actually consume external MCP servers in a way Sam can configure himself (see Sprint 0 spike 4), or (b) the tool surface doesn’t match how Sam thinks about allocation. Either way, we ship a thin Excel add-in OR an Astro /allocation panel in the next sprint. Adoption test, not engagement test.
Cross-Model Perspective
Section titled “Cross-Model Perspective”Codex (GPT-5 / 5.4) was unavailable this session — the user’s ChatGPT account plan does not provide either model to the Codex CLI. A Claude subagent ran the cold read instead.
Honest about the lineage of the MCP decision: the subagent and the would-be-Codex were briefed on a panel-architecture framing. They did NOT see the MCP+skills reframe. David made that strategic call in Phase 4 from his own read of where AI infrastructure is going, NOT from cross-model validation. The cold read endorsed sequencing (Mac email first, Sam allocation second, Min’s email untouched) and the “evidence packet not judgment engine” discipline — both of which carry over to the MCP architecture. The cold read did not endorse, evaluate, or see the MCP-only surface choice.
What the subagent contributed that materially shaped this design:
- The Sam quote as product thesis. “…automating that, it is still manual because you kind of have to go look at one who’s bitching the most” — surfaced by the subagent as the load-bearing sentence in the transcript. This becomes the architecture’s discipline: evidence assembly is the product; judgment is the customer’s.
- P2 reframe. Original P2 was “Mac is the buyer, build what he asks for.” Subagent reframed to “the customer email itself is the forwardable artifact, dashboard screenshot is not.” That reframe survived translation to the MCP architecture because the artifact (customer email) is still the same.
- Real-world evidence test for P2. “Ask Mac directly: if Sam left tomorrow, what breaks?” If the answer is “we’d be fine for 2 weeks” then sequencing Mac first is right. If Mac flinches, allocation is co-equal priority. This is included in The Assignment.
- Cuts. Subagent independently picked the same things to cut that the user picked: short-supply email pipeline (Min owns it), auto-allocation logic, anything past Q2C stage 2.
What the subagent missed / got wrong:
- Recommended an Astro
/allocationpanel as the right week-2 build. David rejected that direction in favor of MCP-only. The subagent’s underlying constraint (evidence-packet, no judgment-engine) survives the surface change. - Did not surface P7 (server-side trigger requirement). That gap was visible only after the MCP reframe.
Approaches Considered
Section titled “Approaches Considered”Approach A — Mac’s email only (Minimum Viable, panel-architecture world)
Section titled “Approach A — Mac’s email only (Minimum Viable, panel-architecture world)”Summary: Just ship order-event emails. CK-branded. Two emails per order. Defer everything else.
Effort: S (~10 part-time hours, ≈1 week)
Risk: Low
Pros: Smallest possible ship; renewal artifact lands in customer inboxes; forces an event model to exist.
Cons: Sam still burns 40 hrs/week; no second harness surface; CK doesn’t perceive “platform” expansion.
Reuses: apps/internal/src/lib/email/client.ts, existing order-ingest events, org/customer auth.
Status: Considered, rejected in favor of D (only addresses 1 of 4 pains).
Approach B — Mac’s email + Sam’s allocation panel (Astro, subagent’s recommendation)
Section titled “Approach B — Mac’s email + Sam’s allocation panel (Astro, subagent’s recommendation)”Summary: Week 1 email, week 2 a read-only /allocation Astro panel showing per-SKU available, total ordered, per-customer recent fill rate, days-since-last-short, open POs. Sam decides; no auto-allocation.
Effort: M (~20 part-time hours over 2 weeks, contingent on Fabio’s data pipeline work and an inventory-pipeline spike)
Risk: Medium
Pros: Both forwardable artifact (email) AND churn-prevention artifact (allocation panel) in one sprint; two harness panels live; intact narrative.
Cons: Locks Ordermatic into the “we build a panel for every pain” pattern — exactly the thing David wants to escape; doesn’t position for the AI-client-as-frontend market motion.
Reuses: Revenue Leaks chrome, harness auth/org, Dagster+dlt WhereFour pipeline, email client.
Status: Considered, rejected in favor of D.
Approach C — Excel-as-frontend (Creative/Lateral)
Section titled “Approach C — Excel-as-frontend (Creative/Lateral)”Summary: Email week 1. For allocation, pipe Ordermatic-computed columns BACK into Sam’s existing Excel sheet via Google Sheets / Excel Online API. Meet Sam where he works.
Effort: M (~2 weeks; sheet conflict/auth edge cases)
Risk: Medium-High (sheet write conflicts; Sam loses trust on one bad merge)
Pros: Tests “augment operator’s existing surface” thesis; zero behavior change for Sam.
Cons: No Ordermatic surface visible to Mac; doesn’t produce a future-prospect-facing demo artifact; architectural drift from /revenue-leaks.
Reuses: Email client, warehouse data layer.
Status: Considered, rejected in favor of D.
Approach D — MCP server + skills + auth (no frontend) — CHOSEN
Section titled “Approach D — MCP server + skills + auth (no frontend) — CHOSEN”Summary: Build an authenticated Ordermatic MCP server exposing tools for inventory, orders, customer signals, revenue-leak state, and email actions. Sam (and Min, and any future CK-shaped customer) connects their AI client of choice (Claude Desktop in Sprint 1; ChatGPT connectors / Cursor deferred to Sprint 2+) to the MCP server with an org-scoped token and queries via natural language. Server-side Temporal workflows handle event-triggered actions (order received → email send). No new UI surface. Skill files are deferred per spec-review scope cut — MCP tool descriptions carry the workflow. Effort: M — ~24 part-time hours total: Sprint -1 (~30 min — two Slack messages, see Assignment) + Sprint 0 (~4 hrs spikes, includes Anthropic-Claude-for-Excel-MCP consumption spike) + Sprint 1 (~10 hrs base, +2 days contingency if auth-scope AND Temporal class are both net-new) + Sprint 2 (~10 hrs). Calendar: 2-3 weeks at David’s part-time cadence. No Excel add-in scope under any outcome — the MCP server is the only artifact we ship; Anthropic’s Claude surfaces (Excel plugin or Desktop) are the consumption layer. Risk: Medium (MCP server hosting + auth model is new; Sam’s AI-client adoption is unproven for this user) Pros:
- Strategic alignment with where AI infrastructure is going. Anthropic MCP is a real standard now; ChatGPT has connectors; Cursor uses MCP. Building MCP-first positions Ordermatic as the auth+audit+data layer behind the customer’s AI agents.
- Sharper rebuttal to “we’ll build our own AI agent” cold-call objection (per Adrienne thread in epic-roentgen design): “Use your own AI agent. Just use Ordermatic as the auth-protected, audited, ERP-integrated data layer. Your IT team gets to keep the agent; we handle the parts they don’t want to maintain (auth, write-path, idempotency, multi-customer pattern data, liability transfer).”
- Capability-first surface: each new pain point = new MCP tool + new skill, not a new page. Lower marginal cost per capability.
- Multi-customer pattern data moat compounds inside the MCP tools (they see signals across all orgs), not inside any one frontend.
- Reuses harness v1’s auth/org model and Revenue Leaks data layer for backing storage. Revenue Leaks Astro page is the LAST panel; everything new is MCP-native.
- The “customer email itself” stays the forwardable artifact Mac shows off (P2 unchanged). Cons:
- Sam’s AI-client adoption is unproven. If Sam doesn’t already use Claude Desktop / ChatGPT for ops work, the allocation tools fire zero times. Open Q in this doc.
- Mac sees no Ordermatic UI. The “Ordermatic shipped a platform” perception relies entirely on customer emails being recognizably ours.
- No demo screenshot for next prospect. We’d demo by SHOWING a Claude conversation that called Ordermatic tools — different demo motion, less proven.
- Template editing UX without a frontend is awkward. Defaults to markdown skill files in org config; if Sam/Mac need to tune copy daily, this becomes friction. Reuses: existing email client, existing org/auth, Dagster+dlt WhereFour pipeline (still source of truth for inventory/orders), Revenue Leaks DB schema (extends rather than replaces).
Recommended Approach — D, with explicit scope
Section titled “Recommended Approach — D, with explicit scope”Frontend exception, stated up front: the existing Revenue Leaks Astro page (shipped as ENG-195) is and remains the ONLY Ordermatic frontend surface for CK. Everything in this design is MCP-native. No new Astro pages, no panel additions to the existing layout, no admin shell. A future engineer reading this doc should treat “frontend = Revenue Leaks only” as a load-bearing constraint, not a default to override.
Time-budget convention: all part-time hour estimates use 1 sprint ≈ 10 part-time hours ≈ 1 calendar week. “Sprint 1 ≈ 10 hours” and “Sprint 1 ≈ 1 week” mean the same thing — David’s part-time cadence.
Build the Ordermatic ops MCP server: apps/ops-mcp — Hono-based HTTP transport, copies the scaffold + Dockerfile + deploy pattern from the existing apps/graph-mcp (which is the in-repo reference implementation; separate codebase, same shape). Architecturally distinct domains: graph-mcp exposes codebase-navigation tools to an AI; ops-mcp exposes customer ops data and writes (allocation decisions, email sends) to an AI. No code sharing in v1 except common patterns.
Architecture Diagrams
Section titled “Architecture Diagrams”Order-confirmation email pipeline
Section titled “Order-confirmation email pipeline”extracted_order.received wherefour.order_pushed.succeeded │ │ ▼ ▼┌────────────────────────────┐ ┌─────────────────────────────────┐│ OrderConfirmationEmail │ │ OrderConfirmationEmail ││ Workflow (Temporal) │ │ Workflow (Temporal) ││ stage: "received" │ │ stage: "processed" │└──────────┬─────────────────┘ └────────────┬────────────────────┘ │ │ ▼ ▼┌──────────────────────────────────────────────────────────────────┐│ apps/ops-mcp tool: send_order_*_email(order_id) ││ 1. Idempotency check (has this stage already fired? skip) ││ 2. Resolve customer email via 3-step chain: ││ a) WhereFour customer.external_email (per Sprint 0 spike 2) ││ b) customer_email_overrides (CK org table) ││ c) catch-all (notifications@cleveland-kitchen.com) ││ 3. Render template (markdown → sanitized HTML) ││ 4. Send via apps/internal/src/lib/email/client.ts ││ 5. Insert audit_logs row (eventCategory='mcp', tool=...) ││ 6. Insert email_events row (for the existing visibility table) │└──────────┬────────────────────────────────────────┬──────────────┘ │ success │ failure (retry 1m/5m/30m × 3) ▼ ▼┌────────────────┐ ┌───────────────────────┐│ Customer mail │ │ DLQ (email_send_dlq) ││ delivered │ │ + BetterStack alert │└────────────────┘ │ if depth > 5 in 1hr │ └───────────────────────┘MCP tool call lifecycle
Section titled “MCP tool call lifecycle”Sam in Claude (Anthropic Claude-for-Excel plugin OR Claude Desktop) │ │ POST /mcp/tools/inventory_snapshot │ Authorization: Bearer <opaque-token> ▼┌─────────────────────────────────────────────┐│ apps/ops-mcp (Hono on Cloud Run) ││ ││ 1. Auth middleware ││ ├─ Look up token in mcp_tokens table ││ ├─ Verify not expired/revoked ││ ├─ Load org_id + scopes from row ││ └─ Update lastUsedAt ││ ││ 2. Scope check ││ └─ token has 'mcp:inventory.read'? ││ else 403 ││ ││ 3. Tool implementation ││ └─ Calls wherefour-service.ts or ││ direct DB query, scoped to org_id ││ ALWAYS from token, NEVER from args ││ ││ 4. Response (JSON) ││ ││ 5. Audit log insert ││ audit_logs(org_id, actor=token_id, ││ eventCategory='mcp', ││ eventType='tool_called', ││ metadata={tool, args_hash, latency_ms}) │└──────────┬──────────────────────────────────┘ │ ▼ Sam's Claude receives JSON response. In Excel-plugin path: Claude writes evidence cells back into the active sheet. In Desktop path: Claude renders in chat.Multi-tenant isolation guard (P6 + cross-org test)
Section titled “Multi-tenant isolation guard (P6 + cross-org test)”Token claim (from mcp_tokens row) org_id = "org_CK_clerk_id" scopes = ["mcp:inventory.read", "mcp:allocation.write"] │ ▼┌──────────────────────────────────────────────┐│ Tool handler: customer_fill_rate( ││ customer_id: "cust_789", ◄── from args ││ sku?: "PRO-001", ││ weeks: 8 ││ ) ││ ││ db.select() ││ .from(orderShipments) ││ .where(and( ││ eq(orderShipments.organizationId, ││ TOKEN.org_id), ◄── ALWAYS from ││ token, NEVER ││ from args ││ eq(orderShipments.customerId, ││ args.customer_id), ││ gte(orderShipments.shippedAt, ││ sub(now, weeks=args.weeks)) ││ )) │└──────────────────────────────────────────────┘
GUARANTEE: a CK token attempting to query customer_idthat belongs to Genfit returns an empty result set(zero rows match the org_id + customer_id pair),not a 403. 403 only fires at scope check (step 2 inthe lifecycle diagram above), not at row scope.
Integration test required (Sprint 1): apps/ops-mcp/tests/cross-org-isolation.test.ts - Issue CK-scoped token - Request customer_fill_rate(<Genfit_customer_id>) - Assert: zero rows in response - Assert: audit_logs row recorded the callSprint -1 (Sprint -1 = today, ~20 minutes of David’s time, gates all engineering)
Section titled “Sprint -1 (Sprint -1 = today, ~20 minutes of David’s time, gates all engineering)”Two short conversations before engineering starts. Both gate Sprint 0; an unfavorable answer to either reshapes the architecture.
- Slack Mac: “If Sam left tomorrow, what breaks? For how long?” (Assignment item #1.) Answer shifts Mac-first vs allocation-first sequencing.
- Slack/call Sam (CONFIRMATORY only — David already has the answer): “You all use Claude in Excel for allocation work today — is that the surface you want Ordermatic plugged into, or is there a different Claude surface you actually prefer?” (Assignment item #2.) This is no longer a discovery question (per P8 — CK uses Claude). It’s locking the exact surface so Sprint 0’s MCP-consumption spike targets the right client.
If either answer hasn’t landed by start of Sprint 0, do not start the engineering spike — the Sprint 0 spike budget is at risk from an architecture pivot the conversations could have prevented.
Sprint 0 (Day 1 spikes, ~4 hours total, blocking before tool/workflow code)
Section titled “Sprint 0 (Day 1 spikes, ~4 hours total, blocking before tool/workflow code)”Four load-bearing unknowns get resolved before Sprint 1 implementation:
- MCP SDK + Hono transport spike (~1 hr): Validate
@modelcontextprotocol/sdkHTTP transport works inside a Hono route on Cloud Run. The SDK is TypeScript-first but has no first-party Hono adapter; a thin HTTP-transport wrapper is the expected path. Fallback: dedicated Node server (Express or vanillahttp) if Hono integration is more than half a day. - Customer email mapping spike (~1 hr): For SPS Commerce / EDI 850 orders, the BT/ST segments yield buyer/ship-to account codes, not email addresses directly. Confirm whether a
customer_account_code → emailmapping table exists in WhereFour customer records, a CK-maintained extension, or anywhere else. If <50% mappable from existing data, see remediation branch below. For direct email orders, use the sender address. - Auth scope claim format (~1 hr): Confirm whether
apps/graph-mcp’s existing bearer-token format carries ascopesclaim. If yes, reuse. If no, the auth extension (JWT signed with existing org-auth key; scope claim is a space-separated string matchingmcp:<domain>.<verb>) adds ~2 days to Sprint 1 budget — call out before coding. - Anthropic Claude-for-Excel MCP consumption spike (~1 hr, NEW, load-bearing for P8): Determine how Anthropic’s Claude-for-Excel plugin (which CK already uses) consumes an external authenticated MCP server. Three possible outcomes shape Sprint 1, and none of them require us to ship anything Excel-specific:
- (a) Claude-for-Excel speaks MCP directly to a remote HTTP server with bearer auth — best case. Sprint 1 ships the MCP server; CK configures the existing Claude-for-Excel plugin to point at it; Sam stays in his sheet, Claude renders evidence into adjacent cells. The “blow their socks off” UX.
- (b) Claude-for-Excel can’t yet consume external MCP servers — fallback A. Sam uses Claude Desktop with our MCP configured there; he keeps his existing Excel sheet open alongside; pastes between the two. UX is clunkier but works.
- (c) Sam prefers Cowork / Claude Desktop directly + generated Excel artifacts — fallback B. Sam asks Claude (in Cowork or Desktop) for allocation advice; Claude calls MCP tools; Claude generates an Excel file as artifact with the allocation suggestions. Sam opens or imports it. No live in-Excel surface needed; the durable artifact is the spreadsheet Claude wrote. This path is the architectural reason we don’t need to build an Excel add-in under any outcome: Claude is the Excel-writing engine; we’re the data-substrate engine.
- We do NOT build an Excel add-in (Office.js etc.) under any outcome. Justified by (c): if Anthropic’s surfaces can’t live IN Excel, they can still GENERATE Excel. We ship the data; they ship the artifact. That separation is the MCP-substrate bet’s load-bearing simplification.
Sprint 1 budget math (resolved here to avoid double-counting elsewhere):
- Base Sprint 1 = ~10 hours (assumes scope claim exists AND Temporal class is extensible).
- If Sprint 0 spike 3 shows scope claim is net-new → add 2 days (~6 hrs).
- If Sprint 1 finds the Temporal workflow class for outbound email is net-new (likely true based on commit history) → add 1-2 days (~3-6 hrs).
- Worst case (both net-new): Sprint 1 = ~14-22 hrs. Ship slips by one part-time cadence (~3-4 calendar days). NOT a Sprint 2 blocker — Sprint 2 starts when Sprint 1 ships, not on calendar.
Remediation branch if Sprint 0 spike 2 fails: if <50% of CK’s last-30-days orders are mappable to a deliverable customer email:
- Defer outbound order emails to Sprint 1.5 — same ~10-hour budget as Sprint 1, shifted by the time it takes CK to build the email master (estimated 1-2 weeks elapsed, NOT 10 additional engineering hours).
- Open a CK ops workstream: CK builds the
account_code → emailmaster in WhereFour or a CSV Ordermatic syncs. Mac owns this (per his “Mac is on my ass about emails” stated urgency). - Sprint 1 redirects to MCP-tools-only (audit, revenue leak parity, allocation read primitives) so the sprint still ships something.
- Sprint 1 success gate is rewritten in this branch: “Mac forwards email in 5 days” does NOT apply (no emails ship in Sprint 1). Replacement gate: Mac confirms via Slack within 3 days that CK ops is actively building the email master. If Mac doesn’t engage, the design is wrong about who owns customer-email-mapping ownership — escalate to David.
Sprint 1 (week 1, target ~10 part-time hours, gated on Sprint 0 spikes) — Order events + email automation
Section titled “Sprint 1 (week 1, target ~10 part-time hours, gated on Sprint 0 spikes) — Order events + email automation”MCP tools (read-only + send):
send_order_received_email(order_id)— idempotent; templates from CK org configsend_order_processed_email(order_id, wherefour_order_id)get_order_email_log(order_id)— audit visibility, recently-sent emails per orderlist_failed_emails(limit?, since?)— visibility into the DLQ for the org; lets Sam/Min surface stuck emails via their AI client without a consoleget_email_template(template_id)/update_email_template(template_id, body)— markdown templates stored in DB, org-scoped, audit-logged
Emergency-off feature flag for customer emails (resolved during CEO review):
- New env var
ORDERMATIC_OPS_MCP_EMAIL_SEND_ENABLEDchecked by the Temporal email-send activity before each send. False → activity returns immediately withreason='globally_disabled'logged to audit_logs. - Toggled via Infisical; effective on next Temporal task pick-up (~seconds).
- Sprint 1 scope addition: ~15 min. Closes the “wrong-customer email scales out at 9pm Friday with David at day-job” incident class.
Email design spec (resolved during design review):
Information hierarchy (mobile-first, top-to-bottom):
Order Received template:
- Subject line:
"We got your order — {customer_po_number}" - CK wordmark (32px height, brand black)
- Big number: customer’s PO number (typeset as heading; primary identifier)
- Status line: “Received — being entered into our system”
- Order verification: line item count + ship-to address (one line each)
- What happens next: “You will receive a confirmation email once we have entered your order, typically within 1 business day.”
- Signature: signoff from CK ops team (exact line TBD per CK input)
- Plain-text footer with reply-to + standard compliance
Order Processed template:
- Subject:
"Your order is confirmed — {customer_po} → {wherefour_order}" - CK wordmark
- Heading: “Order Confirmed” + both order identifiers
- Highlighted: estimated ship date
- Line item summary: what’s shipping (gracefully handled even when full)
- Contact for changes
- Signature
- Footer
Order Short Notice template (NEW per design review D3 — added to Sprint 1 scope):
- Subject:
"Order confirmed with partial allocation — {customer_po}" - CK wordmark
- Heading: “Order Confirmed — partial allocation”
- Highlighted: which lines were fully fulfilled vs short, with qty per line (“We were able to fulfill: 4 of 5 cases of pickle red onions. Remaining 1 case will follow on the next shipment OR has been canceled per your PO terms.”)
- Estimated ship date (for fulfilled portion)
- Contact for changes
- Signature
- Footer
This adds one Temporal workflow branch (on wherefour.order_pushed.succeeded, check if any line ratio shipped<ordered → fire Short Notice instead of Processed). +30 min Sprint 1 scope.
Anti-slop blacklist (templates MUST NOT use):
- Purple/violet/indigo gradient backgrounds
- Icons in colored circles as decoration
- 3-column feature grids
text-align: centeron body content- Decorative blob/wave SVGs
- Emoji as design elements (no ”🎉 Your order is in!”)
system-uior-apple-systemas primary display font — pick a real typeface- Hero-photo-with-thin-text-overlay treatments
- Marketing hero copy (“Welcome to Cleveland Kitchen!”, “Your all-in-one ordering solution”)
- Blue-to-purple color schemes
PR template review gate: any new email template that includes the above fails review and gets rejected.
Mobile + accessibility standards (baked in):
- Body text minimum 16px
- All text/background contrast ratio minimum 4.5:1
- Mobile-first single-column layout; max width 600px on desktop
- All embedded images (logo, etc.) require alt text
- Plain-text alternative rendered for every HTML email (most modern email clients show it as fallback; some users prefer it)
- All actionable text (PO#, ship date) selectable for copy-paste, NOT inside images
Subject line patterns:
- Always include the customer’s PO number when known
- Order Received uses “We got your order — {PO}”
- Order Processed uses “Your order is confirmed — {PO} → {our_order_ref}”
- Order Short Notice uses “Order confirmed with partial allocation — {PO}”
- Patterns are scannable in inbox without opening
MCP tool response shape philosophy (resolved during design review):
- Every analytical tool returns the raw number AND interpretation hints (above/below average, percentile band, comparable-peer-count, reference range). Closes the “Sam sees ‘fill rate 0.87’ instead of ‘87% above your customer-base average’” gap.
- Example:
customer_fill_rate(customer_id, sku?, weeks=8)returns:{"customer": { "id": "cust_acme", "name": "Acme Foods" },"fill_rate": {"value_pct": 87,"sample_size": 32,"interpretation": "above-average","reference_band": "75-90% typical for this customer-base"},"recent_shorts_in_window": 2} - Interpretation logic lives server-side in the tool implementation. One source of truth per metric. Sprint 2 includes interpretation taxonomies for: fill_rate (above/typical/below), recent_shorts (none/some/elevated), inventory_snapshot freshness (just-updated/stale-by-hours).
- Primitive lookup tools (
inventory_snapshot,open_orders_for_sku) return structured data without interpretation hints since they’re descriptive not analytical.
Mac visibility via CSR inbox BCC (resolved during eng review):
- Every order-confirmation email (both “received” and “processed”) BCCs the CK CSR inbox (e.g.
csr@cleveland-kitchen.com— confirm exact address with Sam in Sprint -1). - Rationale: Mac uses the CSR inbox as a normal communications hub, so emails land where he already looks; no surveillance feel. Solves the “Mac sees no Ordermatic UI” gap surfaced by the eng-review outside voice. BCC, not CC, so customers don’t see internal threads.
- Configurable per org via the
customer_email_overridesschema OR a newmcp_email_settingsrow (Sprint 1 implementation detail).
Sprint -1 operational hygiene (added during eng review):
- 48-hour cap on Mac’s Slack response to the “if Sam left tomorrow” question; if no reply, default to Mac-first sequencing as planned and proceed. Don’t stall ~0 engineering hours on a 5-minute message that may not come.
Sending domain — pragmatic week-1 default:
- Initial send: from
notifications@ordermatic.cowith friendly-from"Cleveland Kitchen Orders <notifications@ordermatic.co>"andReply-To: orders@cleveland-kitchen.com(or CK’s preferred reply target). - Parallel workstream: kick off CK SPF / DKIM / DMARC DNS conversation the same day Sprint 1 starts, target migration to
notifications@cleveland-kitchen.com(or similar CK-owned subdomain) in week 3-4. This unblocks the Sprint 1 demo without a 1-3 week DNS dependency. - David sends DNS records to whoever manages CK’s domain; Ordermatic doesn’t block on it.
Server-side workflows (P7 load-bearing):
- New Temporal workflow class
OrderConfirmationEmailWorkflowfollowing theapps/webapp/src/temporal/workflows/pattern (substrate is established pererp-submission.ts,auth-enforcement-monitor.ts; the workflow CLASS is net-new but scaffolding is ~½ day, not 1-2 days). - Workflow listens to
extracted_order.receivedevent → callssend_order_received_emailMCP tool internally → logs sent email. - Workflow listens to
wherefour.order_pushed.succeededevent → callssend_order_processed_email. - Idempotency (resolved during eng review — belt-and-suspenders):
- Workflow ID format:
order-confirmation-{order_id}-{stage}where stage ∈ {received,processed}. Temporal natively rejects duplicate workflow-id starts → first layer of defense. - DB unique constraint on
email_events(order_id, stage)→ second layer catches any out-of-band send attempt (manual MCP call, admin script). - Closes the
[async-submit-feedback-loop]learning class.
- Workflow ID format:
Template rendering safety (resolved during eng review):
- Templates stored as markdown in DB (per existing schema decisions above).
- Render pipeline:
markdown-it(orremark) withhtml: false→ DOMPurify → email body OR preview render. - Applied at: (1)
send_order_*_emailMCP tools before email-client send, (2) any admin-page preview that rendersget_email_templateoutput. - Belt-and-suspenders: closes phishing-embed in customer emails AND XSS in any preview surface.
Email-send failure handling:
- Retry policy: exponential backoff at 1m / 5m / 30m (3 retries max).
- DLQ: existing Cloud Tasks DLQ pattern (mirror
apps/email-workerconvention if applicable) OR dedicatedemail_send_dlqtable — Sprint 0 reads existing pattern to decide. - Alerting: BetterStack alert on DLQ depth > 5 in 1 hr → David is paged. (CK / Min are NOT paged — they see failures via the
list_failed_emailsMCP tool. Perfeedback_customer_copy_is_internal.mdDavid owns customer-facing failure narrative, not the customer.) - SLA target for visibility: any email failure visible via
list_failed_emailswithin 60 seconds; alert fires within 5 minutes of DLQ depth threshold.
Auth layer (resolved during eng review — new mcp_tokens table, NOT Clerk M2M, NOT signed-JWT):
- New Drizzle schema
packages/db/src/schema/mcp-tokens.ts:mcp_tokens (id uuid primary key,token_hash text unique not null, -- sha256 of opaque tokenorg_id text not null, -- Clerk organizationIdsubject text not null, -- 'sam' / 'min' / 'workflow:order-confirm'scopes text[] not null, -- ['mcp:inventory.read', ...]created_by text not null, -- Clerk user id of issuercreated_at timestamp default now(),expires_at timestamp, -- nullable; null = no expirylast_used_at timestamp,revoked_at timestamp) - Tokens are opaque strings (
mcp_live_<base64-random-32>); we store sha256(token) intoken_hashand require the original on every request. - Sprint 1 ships a minimal token-issuance endpoint at
/api/admin/mcp-tokens(org-admin scoped via existing Clerk session check) — issue / list / revoke. No CLI yet. - Per-tool scopes:
mcp:<domain>.<verb>, e.g.mcp:email.send,mcp:email.read,mcp:inventory.read,mcp:allocation.write,mcp:revenue_leaks.read,mcp:audit.read. - Token issuance for CK in Sprint 1: one per stakeholder (Sam, Min) + one for the Temporal workflow service identity. Mac: no token.
- Min explicitly gets:
mcp:email.read,mcp:inventory.read,mcp:revenue_leaks.read(andmcp:revenue_leaks.writeif we add it in week 3). He does NOT getmcp:email.sendin v1 — his existing short-supply pipeline keeps owning sends per P5. If Min asks for centralization later, that’s the conversation to revisit it. - Sam explicitly gets: all read scopes,
mcp:allocation.writeonce Sprint 2 ships, NOmcp:email.sendin v1 (email sends are server-triggered). - Mac: no token in v1. He never queries the system directly.
Audit log (resolved during eng review — REUSES existing audit_logs table, NOT new; partitioning added):
- The existing
packages/db/src/schema/audit-logs.tstable is Clerk-org-scoped, has JSONBmetadata, indexed byorganizationId + eventType + createdAt. We add new rows with:eventCategory = 'mcp'eventType = 'tool_called'(or'token_issued'/'token_revoked'for admin actions)actorId = mcp_token.subject(e.g.'sam', or'workflow:order-confirm')targetType = 'mcp_tool',targetId = tool_namemetadata = {tool, args_hash, result_status, latency_ms}(args_hash, not raw args, to avoid PII)
- Retention v1 (revised after outside-voice review, ~30 min Sprint 1): single composite index
(organization_id, event_category, created_at DESC)+ a weekly cron DELETE for rows older than 13 months. Saves ~2.5 hrs vs partitioning; sufficient for customer-#2 volume (estimated <100K rows/month across all orgs). - Retention v2 tripwire (documented growth path): when
audit_logsrow count exceeds 5M OR p95 query latency on the table exceeds 200ms, escalate to monthlyPARTITION BY RANGE (created_at)with scheduled archive-and-drop. Same growth-path-documented pattern as the analytics-query BigQuery tripwire (500ms p95). - Customer-facing pull:
get_audit_log(since, limit)MCP tool, org-scoped. Lets CK pull their own audit trail without filing a support ticket. Filters toeventCategory='mcp'so it doesn’t surface unrelated org events.
Postgres connection pooling for ops-mcp on Cloud Run:
- Mirror the pattern apps/webapp uses today (Sprint 0 spike 3 confirms; likely Neon’s serverless driver with connection multiplexing OR PgBouncer fronting CloudSQL). ops-mcp does NOT open raw
pgpools per Cloud Run instance — that exhausts the Postgres connection limit at cold-start scale.
Data isolation (multi-tenancy, load-bearing for P6 future-customer story):
- EVERY MCP tool implementation MUST derive
org_idfrom the token claim, NEVER from a tool argument. A CK token attempting to fetch a Genfit-owned record returns 403, even if the tool signature accepts acustomer_idthat happens to match a Genfit customer. - Sprint 1 ships with an integration test asserting cross-org access is rejected at the tool layer. Add
apps/ops-mcp/tests/cross-org-isolation.test.tsto the test plan.
Customer email mapping path (Sprint 0 spike 2 outcome shapes this):
- BT/ST segments yield account codes. Account codes resolve via:
- First: WhereFour customer record
external_emailfield if populated (TBD in Sprint 0) - Second: CK org’s
customer_email_overridestable in our DB (CSV-imported during onboarding; CK ops can extend via MCPupdate_customer_email(account_code, email)) - Third:
notifications@cleveland-kitchen.com(catch-all fallback Mac forwards manually)
- First: WhereFour customer record
- Per-customer suppression: row in DB, defaults to enabled (sends ON). Toggled via
set_customer_email_preference(account_code, enabled)MCP tool.
Added to Sprint 1 scope during eng review:
.github/workflows/deploy-ops-mcp.yml(copy + adaptdeploy-webapp.yml/deploy-temporal.ymlpattern; ~30 min)- Update
.github/workflows/deploy-branch-gate.ymladdingapps/ops-mcp/**path so PRs intogcp-staging/gcp-prodrun scoped tsc + Dockerfile build (closes the documented hallucinated-import failure mode per learning[validate-yml-misses-gcp-staging-branch]). - New
packages/db/src/schema/mcp-tokens.tsmigration +customer_email_overridesmigration (Sprint 0 spike 2 confirms the override-table shape). - Minimal admin endpoint
/api/admin/mcp-tokens(issue / list / revoke).
Test coverage (resolved during eng review — Sprint 1 must include):
- Cross-org isolation integration test (
apps/ops-mcp/tests/cross-org-isolation.test.ts, 3 scenarios): CK token + CK customer → rows returned; CK token + Genfit customer → zero rows (not 403); audit_logs row recorded. Non-negotiable — load-bearing for P6. - Auth middleware test suite (
apps/ops-mcp/tests/auth-middleware.test.ts, 5 scenarios): valid + scope → 200 + lastUsedAt; unknown token → 401; expired → 401; revoked → 401; valid + missing scope → 403. - Customer email mapping resolution chain (
apps/ops-mcp/tests/customer-email-resolution.test.ts, 5 scenarios): each of the 3 fallback tiers; suppression-toggle-off → no send; ambiguous tier 1 → audit warning + fallback. - Temporal workflow (
apps/webapp/src/temporal/workflows/order-confirmation-email.test.ts, 5 scenarios mirroringerp-submission.test.ts): received-stage start; duplicate workflow-id rejected; activity transient failure → retry; activity hard failure → DLQ + workflow complete-with-error; processed-stage on event. Uses TemporalTestEnvironmentfor time control. - DLQ → BetterStack alert (
apps/ops-mcp/tests/dlq-alert-config.test.ts): asserts the BetterStack alert resource is declared with documented threshold (5 in 1hr). PLUS one-time manual smoke: force 6 DLQ rows, confirm page lands on David. Closes the[email-pipeline-dead-since-april-2026]silent-failure class. - Per-tool happy-path + auth-fail + scope-fail tests — each MCP tool ships with 3 minimum tests in its
.test.tsneighbor file. Baseline coverage that fails the build if missing.
Cut from sprint 1:
- Allocation tools (Sprint 2)
- Any judgment-layer logic
- Real-time email preview (defer to sprint 2 if Mac asks)
- Skill files (subagent review identified these as not load-bearing for v1; MCP tool descriptions are sufficient for an AI client to compose workflows. Skill files reintroduced ONLY if Sprint 2 shows tool descriptions can’t compose well enough to drive Sam’s flow.)
- ChatGPT connectors and Cursor MCP support — Sprint 1 ships Claude-in-Excel + Claude Desktop only (per P8). ChatGPT connectors and Cursor are deferred to Sprint 2+ unless Sam specifies a non-Claude client in Assignment item #2 (unlikely per P8).
Sprint 2 (week 2, target ~10 part-time hours, gated on Sprint 1 ship + Fabio data work) — Allocation evidence packet
Section titled “Sprint 2 (week 2, target ~10 part-time hours, gated on Sprint 1 ship + Fabio data work) — Allocation evidence packet”MCP tools (trimmed per spec-review scope cut — 4 read + 1 write):
inventory_snapshot(sku, location?)— current on-hand, by SKU; sourced from existingdlt_wherefour_inventorywarehouse tableopen_orders_for_sku(sku, window_days?)— total ordered, broken out per customercustomer_fill_rate(customer_id, sku?, weeks=8)— rolling 8-week fill raterecent_shorts(customer_id?, window_days=30)— CONDITIONAL: ships only if warehouse has short-shipment data on Day 1. If not, descope to a follow-up sprint; Sam composes the same view fromcustomer_fill_ratefor now.record_allocation_decision(sku, decisions: [{customer_id, allocated_qty, rationale}])— Sam’s AI logs his decision back into the system; enables next-time suggestion improvement
Explicitly cut from Sprint 2 (per spec review):
customer_priority_signals(customer_id)composite — the AI client composes this from primitives. Premature abstraction at this stage.- Skill files (same reasoning as Sprint 1; the AI client composes the workflow from tool descriptions). Reintroduce if Sprint 2’s first real Sam run shows the AI can’t string the primitives together without help.
Data access pattern (resolved during eng review against the ACTUAL repo state):
Repo’s real data layout:
- Bronze: Iceberg on R2 (per
apps/dagster/drop_iceberg_tables.pyworking undererp_data_native_v2/prefix). - Gold: Parquet on R2 (per
apps/dagster/scripts/direct_typesense_sync.pyheader: “reads Gold Parquet from R2 and pushes to Typesense”). - Typesense: the ops-data query layer for current-state lookups (products, customers, shipping addresses, inventory) —
apps/webapp/src/services/erp-query/typesense-query-service.tsis the established pattern. - Postgres app DB: app state + feature-specific silver materialized tables. The
packages/db/src/schema/revenue-leaks.ts:60revenueLeakSilverLineItemstable is the precedent for analytics-grade silver bridged from Gold via dagster.
ops-mcp tool routing by tool kind:
| Tool kind | Examples | Data path |
|---|---|---|
| Current-state lookups | inventory_snapshot, open_orders_for_sku | Typesense (call into the existing typesense-query-service.ts pattern OR directly to Typesense via its TS client; no apps/webapp import) |
| Time-windowed analytics | customer_fill_rate, recent_shorts | NEW Postgres silver tables Fabio’s dagster pipeline materializes from Gold Parquet. Same pattern as revenueLeakSilverLineItems. Drizzle reads. |
| Writes | record_allocation_decision, send_*_email, audit, mcp_tokens | Postgres app DB via Drizzle |
Analytics performance path (v1 + v1.5 documented):
- v1: Postgres silver tables with compound indexes on
(org_id, customer_id, sku, shipped_at). Sub-10ms p95 at CK’s volume (hundreds of customers, thousands of SKUs). - Every ops-mcp tool logs
latency_mstoaudit_logs.metadatafor telemetry. - Tripwire: any tool exceeds 500ms p95 over 30 days → escalate to v1.5 (BigQuery external tables over Gold Parquet on R2, parallel query path in ops-mcp routing by query complexity). Substrate is already in place; the upgrade is documented, not speculative.
Data foundation (Fabio’s commitment, expanded after eng review — blocks Sprint 2 start):
- Per-tool data target decision: which fields land in Typesense (real-time lookups) vs Postgres silver (analytics windows). Sprint 2 Day 1 task with Fabio.
- New Postgres silver schemas in
packages/db/src/schema/for fill-rate-source and shorts-source tables. Mirrorrevenue_leak_silver_line_itemsshape:clerkOrganizationId+ indexed time column + business keys. - New dagster sync jobs in
apps/dagster/scripts/(mirrordirect_typesense_sync.pyfor Typesense pushes; mirror revenue-leaks Postgres silver sync for the analytics tables). - Sam’s allocation Excel schema feeds at least one of the above targets (likely the Postgres silver side — fill-rate is the analytic shape).
- Sam’s allocation Excel schema → silver/gold table in warehouse. Same
packages/evalpattern as VMS golden datasets per memory. - Inventory-pipeline spike Day 1 (now Sprint 0 alongside the MCP spikes): does
dlt_wherefour_inventoryhave per-location, per-SKU on-hand snapshots? If yes, build on top. If no, extend the dlt source (~2-3 days, eats Sprint 2 budget). - Short-shipment table presence (Day 1 check): determines whether
recent_shortsships in Sprint 2 or gets deferred.
Sprint 2 fallback if Fabio slips past Sprint 1 end:
- Sprint 2 SWAPS allocation MCP tools for
get_revenue_leaks+mark_leak_stateMCP parity work (Open Q #6) so the sprint still ships something MCP-native. - Allocation tools slip to Sprint 3 once Fabio’s data is in. This is a known fallback, not an emergency reroute.
Cut from sprint 2 (intent layer):
- Auto-allocation suggestion logic (P4 — Sam decides)
- “Why this customer ranks higher” automated reasoning (let the AI client surface that from the data; don’t bake into tools)
- Generalization to a second customer (defer to v2.1, same as harness v1 plan)
Sprint 1 + 2 success looks like
Section titled “Sprint 1 + 2 success looks like”Sam’s screen, in his existing Excel allocation sheet: he opens the Claude sidebar in Excel, types “What allocation should I do for pickle red onions next week?”. Claude calls 4-5 Ordermatic MCP tools, assembles inventory + open orders + per-customer fill rates + recent shorts, renders the evidence packet directly into adjacent cells in the sheet, and presents a ranked recommendation in the chat. Sam reviews, edits the quantities in cells where he wants to override, and tells Claude “log it.” Claude calls record_allocation_decision. Sam never opened a browser, never opened a separate app, never copy-pasted anything. The Excel-as-cockpit pattern is the unlock.
In parallel, on the customer-facing side: when a customer’s PO lands in CK’s inbox, Ordermatic’s server sends “we received your order” automatically. Within minutes of CK’s reviewer pushing the order to WhereFour, the server sends “your order has been confirmed in our system.” Both emails are branded for CK, template-edited via MCP from inside Sam’s or Min’s Excel-Claude session.
The two demo videos: (1) Sam in Excel asking Claude what allocation to do, watching evidence appear in cells, deciding. (2) Mac forwarding a customer reply that says “thanks for the order confirmation.”
Open Questions
Section titled “Open Questions”- RESOLVED (per P8 — David confirmed in session): CK uses Claude already, primarily in Excel for ops work. Sprint -1 Assignment item #2 is now confirmatory (“which Claude surface specifically”), not discovery. The Sprint 0 spike that carries the highest variance impact is now spike 4 — does Claude-in-Excel consume external MCP servers natively today, or does Sprint 1 need an Office.js add-in?
- For SPS Commerce orders, where is the end-customer email address? EDI 850 segments? CK’s account-master? Built into WhereFour customer records? Resolution: Sprint 1 Day 1 task — confirm the data path before building
send_order_received_email. - Where does the MCP server run, and how do customers authenticate their AI client to it? Cloud Run with org-scoped bearer tokens is the obvious answer (parity with
apps/graph-mcp). But the AI-client auth UX is new — Claude Desktop / ChatGPT connectors / Cursor each have different MCP auth flows. Resolution: Sprint 1 Day 2-3 scoping task; can default to bearer-token + Claude Desktop config first, others later. - Does Mac actually care that he sees no Ordermatic UI? Possibly. Some buyers want to log in and see a dashboard. Resolution: subagent’s “if Sam left tomorrow, what breaks?” question is the right wedge — same conversation reveals whether Mac feels he has visibility into ops via Sam or via tools.
- Liability rider — same as v1. Still defer until Approach C-flavored H2 thinking.
- Does this design imply Revenue Leaks ALSO gets MCP tools? Probably yes for parity (
get_revenue_leaks,mark_leak_state) — cheap once the server exists. Not in sprint 1 or 2; revisit week 3.
Success Criteria
Section titled “Success Criteria”Sprint 1 gate (end of week 1):
- Three production orders land via SPS Commerce / direct email in CK’s inbox; “received” email goes out for all three within 5 minutes of ingest; “processed” email goes out for all three within 5 minutes of WhereFour push.
- Mac forwards at least one of those emails to David or someone external with positive framing (“look at this” / “thanks for shipping this”) within 5 days of go-live. This is the renewal artifact test.
Sprint 2 gate (end of week 2, conditional on Sprint 0 spike 4 outcome):
- Only applies if Sprint 0 spike 4 lands in outcome (a) or (b) — Claude-in-Excel or Claude Desktop can consume the MCP server. If outcome (c) fires, Sprint 1 absorbed the Office.js add-in scope; this Sprint 2 gate shifts to Sprint 3.
- Sam connects Claude-in-Excel to the MCP server using the org-scoped token David issues him.
- Sam runs at least one full allocation cycle FROM HIS EXISTING ALLOCATION SHEET via Claude-in-Excel + MCP tools within 7 days of go-live.
- That cycle takes less than 8 hours of Sam’s time (vs his current ~40 hours).
- The visible demo artifact: a screen recording of Sam’s allocation Excel sheet with Claude rendering evidence into cells and Sam deciding. This is the prospect-facing demo that replaces “log into our dashboard.”
Two-week post-launch gate:
- Mac has forwarded 3+ customer emails OR comments positively to David about them within 14 days.
- Sam has called MCP tools from his AI client across 3+ different days in the 14-day window.
Kill criteria (hard, quantitative — no rationalization):
- Email adoption kill: zero emails sent in 14 days post-launch (technical failure, not customer demand). Investigation, not architecture pivot.
- MCP adoption kill: Sam calls zero MCP tools from Claude (Excel or Desktop) in 14 days post-launch AND is still doing his full manual Excel workflow → P8 is wrong (Anthropic’s Claude surfaces + our MCP doesn’t fit his real allocation flow) OR the tool surface doesn’t match how Sam thinks. Pivot decision: (a) ship a thin Astro
/allocationpanel over the same MCP tools, OR (b) revalidate at next signed WhereFour customer. We do NOT build an Excel add-in as a fallback — the MCP-substrate-only bet stands or falls on Anthropic surfaces working for the customer. - Customer-email-mapping kill: less than 50% of orders can be matched to a deliverable customer email address → suppress non-matched, send only matched, escalate the data-completeness gap to CK ops as a separate workstream.
Distribution Plan
Section titled “Distribution Plan”- New service:
apps/ops-mcpdeployed to Cloud Run via existing CI/CD (mirrorapps/graph-mcppipeline). - Existing Astro app unchanged (Revenue Leaks page stays).
- Customer connects via Claude-surface-specific MCP configuration (Sprint 0 spike 4 outcome shapes which one is primary):
- Claude-in-Excel: primary surface for Sam. Configuration path depends on spike outcome — either native MCP config or Office.js add-in that wraps Anthropic API + MCP server-side.
- Claude Desktop: secondary surface for Min and for non-allocation workflows (revenue-leak triage).
claude_desktop_config.jsonsnippet documented in CK’s customer onboarding doc (perfeedback_customer_doc_conventions.mdlayered storage — Coda canonical, gbrain mirror, worktree authoring surface). - ChatGPT connectors / Cursor: deferred to Sprint 2+ (per P8, unlikely CK needs either).
- Same SKU; per-tool scopes documented in the org’s auth config.
- No new release surface for emails; they ship as a Temporal workflow update.
Dependencies
Section titled “Dependencies”- Anthropic MCP SDK (
@modelcontextprotocol/sdk) for the server; HTTP transport adapter inside Hono is a Sprint 0 spike (no first-party adapter exists today). - WhereFour silver-layer line-item granularity (unknown — same Open Question as v1; recheck in Sprint 0).
- Sam’s allocation sheet (Fabio’s data sync) — David committed verbally. Sam needs to share the file before Sprint 2 begins.
- Email deliverability: Sprint 1 sends from
notifications@ordermatic.cowith friendly-from to dodge the 1-3 week DNS conversation; CK SPF / DKIM / DMARC kicks off in parallel for week-3-4 migration to a CK-owned subdomain. - Customer-contact-mapping data path (Open Question 2 / Sprint 0 spike 2) — Sprint 1’s email-send tools work fine with a partial mapping; the remediation branch handles <50% mappability without blocking Sprint 1 entirely.
- Auth scope claim format (Sprint 0 spike 3) — if
apps/graph-mcpalready supports scopes, zero new auth work; if not, add ~2 days to Sprint 1 budget. - New Temporal workflow class for email-send (
OrderConfirmationEmailWorkflow) — assumed net-new based on commit history (no prior outbound-email workflow exists); 1-2 days inside Sprint 1’s 10-hour envelope. - Symphony orchestrator capacity to execute the build per David’s verbal commitment (“I’ll get my factory working on it”).
- David’s part-time engineering: Sprint -1 (~30 min Slack) + Sprint 0 (~3 hrs spikes) + Sprint 1 (~10 hrs base; +2 days contingency if scope claim AND Temporal class are both net-new) + Sprint 2 (~10 hrs) = ~23 hrs base, ~30 hrs worst case, over 2-3 calendar weeks.
The Assignment
Section titled “The Assignment”Three things this week, before any code is written:
1. Send Mac a Slack message with the specific question the subagent surfaced: “If Sam left tomorrow, what breaks? For how long?” The answer tells you whether to keep the Mac-first sequencing or flip to allocation-first. 5 minutes of his time; saves potentially 2 weeks of mis-sequenced engineering.
2. Confirm with Sam which Claude surface he uses for allocation work today. Per P8 you already know CK uses Claude in Excel; this is a confirmation conversation to lock the specific surface (Claude in Excel sidebar? Claude in browser alongside Excel? Both?). The answer shapes Sprint 0 spike 4 — does Claude-in-Excel speak to a remote MCP server natively, or does Sprint 1 add an Office.js add-in? 10 minutes of Sam’s time.
3. Get Sam to share the allocation sheet with Fabio. Verbally committed on the call; this is the literal “do the thing you said you’d do” follow-up. Until the sheet lands, sprint 2 is blocked even if sprint 1 ships.
The engineering build follows these three — all are 5-15 minute conversations and they lock in the design’s accountability before code starts.
Known risks accepted (CEO review)
Section titled “Known risks accepted (CEO review)”- mcp_tokens have no expiry, no rotation, no IP allowlist in v1 (D3). Long-lived bearer credentials on customer machines. If Sam’s or Min’s laptop is compromised, attacker gets org-scoped read+write access to CK data AND can send customer-facing emails as CK. Mitigated only by audit logging and the CSR-inbox BCC (which would surface unusual email volume). Token hardening (TTL, rotation, IP allowlist, new-IP alert) captured as a P1 TODO; revisit before customer #3 or before fundraising due-diligence reveals it as a question.
NOT in scope
Section titled “NOT in scope”Considered and explicitly deferred during eng review:
- Excel add-in (Office.js) — not under any Sprint 0 spike 4 outcome. Claude generates Excel artifacts; we don’t embed in Excel. Architectural commitment.
- ChatGPT connectors / Cursor MCP support — Sprint 1 ships Claude-in-Excel + Claude Desktop only (per P8). Defer until non-Claude client demand emerges.
customer_priority_signalscomposite tool — let Claude compose from primitives (customer_fill_rate,recent_shorts). Premature abstraction at v1.recent_shortstool — CONDITIONAL on Day-1 warehouse readiness check. If short-shipment data isn’t in silver, slips to a follow-up.- Skill files — MCP tool descriptions carry the workflow. Reintroduce only if Sprint 2 shows Claude can’t compose without them.
- BigQuery analytics path — documented v1.5 tripwire at 500ms p95. Boring Postgres now, columnar growth path written down.
audit_logsmonthly partitioning — documented v2 tripwire at 5M rows OR 200ms p95 query latency. Composite index + weekly DELETE for now.- Auto-allocation logic — Sam decides (P4). Evidence-packet not judgment-engine.
- Generalization to a second customer — v2.1 once CK kill-criteria pass.
- Liability rider — Approach C H2 question. Insurer conversation deferred.
- ops-mcp admin UI for token issuance — Sprint 1 ships only the endpoints. UI captured as TODO in TODOS.md.
- Min’s short-supply pipeline integration — P5 says don’t touch. Captured as TODO gated on Min asking.
What already exists (and the design now credits)
Section titled “What already exists (and the design now credits)”Surfaced during the Step-0 scope challenge:
apps/graph-mcp/src/mcp.ts— established MCP server pattern in repo.apps/ops-mcpcopies the scaffold + deploy shape.apps/webapp/src/temporal/workflows/— Temporal substrate is mature (erp-submission,auth-enforcement-monitor,workflow-ownership). NewOrderConfirmationEmailWorkflowis class-net-new but ~½ day scaffold, not 1-2 days.packages/db/src/schema/audit-logs.ts— existing org-scoped audit table. ops-mcp reuses witheventCategory='mcp'(DRY win, no parallel table).apps/webapp/src/services/erp/implementations/wherefour-service.ts— WhereFour data access. ops-mcp doesn’t import this service (decoupling), but the patterns mirror.apps/webapp/src/services/erp-query/typesense-query-service.ts— Typesense is the established ops-data read layer. ops-mcp lookups route here.packages/db/src/schema/revenue-leaks.ts—revenueLeakSilverLineItemsis the precedent for feature-specific Postgres silver. ops-mcp’s allocation analytics tables mirror this pattern.apps/internal/src/lib/email/client.ts— outbound email client. ops-mcp’s email tools call this..github/workflows/deploy-webapp.yml,deploy-temporal.yml— deploy pattern to copy fordeploy-ops-mcp.yml.deploy-branch-gate.ymlis the registration point.
Failure modes per new codepath
Section titled “Failure modes per new codepath”Per skill protocol: each new codepath gets one realistic production failure scenario, with test/error-handling/visibility assessment.
| Codepath | Realistic failure | Test? | Error handling? | User-visible? |
|---|---|---|---|---|
| Auth middleware | mcp_tokens query times out → 500 | T2 (D11) covers happy + error returns | Yes — middleware catches and returns 503 | Yes — Claude shows error |
| Cross-org guard | New tool forgets org_id filter in WHERE | T1 (D10) integration test | None at runtime (silent leak) | NO — would be silent in prod CRITICAL if test ever skipped |
send_order_received_email | Email mapping resolves to wrong contact (shipping vs AP) | T3 (D12) covers tier ambiguity → audit warning | Audit row logs which tier resolved; suppression possible | CRITICAL SILENT MODE — wrong customer receives email; no auto-detection until Mac/customer complains. Mitigation: BCC CSR inbox makes the wrong-recipient visible to CK ops immediately. |
| Temporal workflow | Workflow-id collision (same order, same stage, simultaneous events) | T4 (D13) covers idempotency | Temporal natively rejects 2nd; DB constraint catches anything out-of-band | Yes via list_failed_emails if it bubbles up |
| DLQ → BetterStack | Alert config drifts after a terraform refactor | T5 (D14) integration test verifying config exists | Manual smoke on every config touch | Degraded if both test and smoke skipped |
| Customer email mapping | <50% mappable from existing data | Sprint 0 spike 2 catches; remediation branch documented | Remediation branch: defer email-sending Sprint 1, Mac owns CK-side data master | Yes — surfaced in Sprint -1 conversation |
Critical gaps (failure mode is silent AND lacks one of test/error-handling/visibility):
- Cross-org guard skipped accidentally on a new tool added later — mitigated by the integration test running on every PR, BUT the regression rule fires only if T1’s test pattern is mirrored for every new tool. Action: add lint rule or convention test that every tool handler unit-tests cross-org rejection. Captured implicitly in “Per-tool baseline tests (happy + 401 + 403)” requirement; add a 4th test: “cross-org rejection.”
- Wrong-contact email — silent failure mode mitigated by the CSR inbox BCC pattern (D18). Mac sees the email in his normal inbox and notices “wait, that went to the wrong person.” This is the load-bearing mitigation; without the BCC, this gap is critical.
Worktree parallelization strategy
Section titled “Worktree parallelization strategy”Sprint 1 work splits into 4 lanes; Sprint 2 work splits into 2 lanes.
Sprint 1 dependency table
Section titled “Sprint 1 dependency table”| Step | Modules touched | Depends on |
|---|---|---|
| 1A: ops-mcp scaffold + auth middleware | apps/ops-mcp/ | — |
| 1B: Temporal workflow + email-send activity | apps/webapp/src/temporal/workflows/, apps/webapp/src/temporal/activities/ | 1C (mcp_tokens schema needed for token-auth in workflow’s MCP call) |
| 1C: Schemas (mcp_tokens, customer_email_overrides, email_events unique constraint, audit_logs composite index) | packages/db/src/schema/, packages/db/src/migrations/ | — |
| 1D: CI/CD (deploy-ops-mcp.yml + deploy-branch-gate update) | .github/workflows/ | 1A (need the service to deploy) |
| 1E: Tests (cross-org, auth, customer-email, workflow, DLQ alert) | apps/ops-mcp/tests/, apps/webapp/src/temporal/workflows/*.test.ts | 1A + 1B + 1C |
1F: Admin token endpoints (/api/admin/mcp-tokens) | apps/webapp/src/pages/api/admin/ | 1C |
Sprint 1 lanes
Section titled “Sprint 1 lanes”- Lane A: 1C → 1A → 1F → 1E (sequential, ~6-7 hrs)
- Lane B: 1C → 1B → 1E-workflow-tests (sequential, ~3 hrs)
- Lane D: 1A → 1D (sequential, ~1 hr)
Lane A and Lane B fork after 1C (schemas) lands; both contribute to 1E. Lane D forks after 1A scaffolding lands.
Launch order: 1C first (blocks both A and B). Then A + B + D start in parallel. 1E consolidates at the end.
Sprint 2 dependency table
Section titled “Sprint 2 dependency table”| Step | Modules touched | Depends on |
|---|---|---|
| 2A: Fabio’s dagster pipelines for allocation silver tables | apps/dagster/erp_pipeline/, apps/dagster/scripts/ | Sam’s sheet shared (Sprint -1 deliverable) |
| 2B: Postgres silver schemas for analytics tools | packages/db/src/schema/ | 2A (informs schema shape) |
| 2C: MCP tools (inventory_snapshot, open_orders_for_sku, customer_fill_rate, recent_shorts, record_allocation_decision) | apps/ops-mcp/src/tools/ | 2B for analytics tools; Typesense already exists for lookup tools |
| 2D: Per-tool tests + latency instrumentation | apps/ops-mcp/tests/, apps/ops-mcp/src/middleware/ | 2C |
Sprint 2 lanes
Section titled “Sprint 2 lanes”- Lane E: 2A → 2B → 2C-analytics-tools → 2D
- Lane F: 2C-lookup-tools (Typesense path, independent of Fabio’s work) → 2D
Lane F can ship lookup tools (inventory_snapshot, open_orders_for_sku) on day 1 of Sprint 2 because Typesense already has the data. Lane E lands analytics tools later in the sprint after Fabio’s pipeline materializes.
Conflict flags: none. Lanes touch different module trees.
Implementation Tasks
Section titled “Implementation Tasks”Synthesized from this review’s findings. Each task derives from a specific decision. P1 blocks ship; P2 should land same branch; P3 is a follow-up.
-
T1 (P1, human: ~30min / CC: n/a) — Sprint -1 conversations
- Surfaced by: design, D17, D18 (3 messages: Mac timeout-48h, Sam Claude-surface confirmation, Anthropic docs read)
- Files: none (Slack)
- Verify: 3 answers landed OR 48h timeout fired
-
T2 (P1, human: ~1h / CC: ~15min) — Sprint 0 spikes 1-4
- Surfaced by: Sprint 0 in design
- Files: scratch / notes
- Verify: 4 spike outcomes recorded; Sprint 1 budget adjusted per math
-
T3 (P1, human: ~30min / CC: ~10min) —
apps/ops-mcp/scaffold (copyapps/graph-mcp/pattern)- Surfaced by: A1, D2
- Files:
apps/ops-mcp/{package.json, Dockerfile, src/server.ts, src/index.ts, tsconfig.json} - Verify: ops-mcp starts locally + responds to a hardcoded test tool
-
T4 (P1, human: ~30min / CC: ~10min) —
packages/db/src/schema/mcp-tokens.ts+ migration- Surfaced by: A3, D4
- Files:
packages/db/src/schema/mcp-tokens.ts, migration SQL - Verify: pnpm db:migrate succeeds; insert + select works via Drizzle
-
T5 (P1, human: ~30min / CC: ~10min) —
customer_email_overridesschema + migration- Surfaced by: D17 customer email mapping chain
- Files:
packages/db/src/schema/customer-email-overrides.ts, migration SQL - Verify: insert + select via Drizzle
-
T6 (P1, human: ~15min / CC: ~5min) — email_events unique constraint + audit_logs composite index
- Surfaced by: D9 (idempotency), D19 (audit perf)
- Files: migration SQL
- Verify: pg index hits visible in
EXPLAINagainst composite query
-
T7 (P1, human: ~1.5h / CC: ~30min) — Auth middleware (token lookup + scope check + lastUsedAt + cross-org guard)
- Surfaced by: A3, T1, T2 (D10, D11)
- Files:
apps/ops-mcp/src/middleware/auth.ts, isolation helper - Verify: unit tests pass (T11 below)
-
T8 (P1, human: ~2h / CC: ~30min) —
OrderConfirmationEmailWorkflow(received + processed stages)- Surfaced by: design, CQ4/D9
- Files:
apps/webapp/src/temporal/workflows/order-confirmation-email.ts,apps/webapp/src/temporal/activities/order-confirmation-email-activity.ts - Verify: workflow tests pass (T13 below)
-
T9 (P1, human: ~2h / CC: ~30min) — MCP tools:
send_order_received_email,send_order_processed_email+ email resolution chain + markdown render- Surfaced by: design, CQ3/D8, T3/D12
- Files:
apps/ops-mcp/src/tools/send-order-received.ts,send-order-processed.ts,email-resolution.ts,markdown.ts - Verify: customer-email-resolution test passes (T12)
-
T10 (P1, human: ~1.5h / CC: ~30min) — Remaining Sprint 1 MCP tools (
list_failed_emails,get_audit_log,get_order_email_log,get_email_template,update_email_template,update_customer_email,set_customer_email_preference)- Surfaced by: design Sprint 1 tools list
- Files:
apps/ops-mcp/src/tools/*.ts - Verify: per-tool baseline tests (happy + 401 + 403 + cross-org) pass
-
T11 (P1, human: ~1h / CC: ~20min) —
apps/ops-mcp/tests/cross-org-isolation.test.ts(3 scenarios) +auth-middleware.test.ts(5 scenarios)- Surfaced by: T1/D10, T2/D11
- Files:
apps/ops-mcp/tests/cross-org-isolation.test.ts,apps/ops-mcp/tests/auth-middleware.test.ts - Verify: pnpm test passes
-
T12 (P1, human: ~1h / CC: ~20min) —
apps/ops-mcp/tests/customer-email-resolution.test.ts(5 scenarios)- Surfaced by: T3/D12
- Files:
apps/ops-mcp/tests/customer-email-resolution.test.ts - Verify: pnpm test passes
-
T13 (P1, human: ~1.5h / CC: ~30min) —
apps/webapp/src/temporal/workflows/order-confirmation-email.test.ts(5 scenarios using TestEnvironment)- Surfaced by: T4/D13
- Files:
apps/webapp/src/temporal/workflows/order-confirmation-email.test.ts - Verify: pnpm test passes including retry timing
-
T14 (P1, human: ~30min / CC: ~10min) —
apps/ops-mcp/tests/dlq-alert-config.test.ts+ one-time manual smoke- Surfaced by: T5/D14
- Files:
apps/ops-mcp/tests/dlq-alert-config.test.ts - Verify: terraform/Infisical declaration assertion passes; manual force 6 DLQ rows → BetterStack page lands
-
T15 (P1, human: ~30min / CC: ~10min) —
/api/admin/mcp-tokensendpoints (issue/list/revoke, Clerk-session gated)- Surfaced by: A3/D4 (Sprint 1 admin endpoint requirement)
- Files:
apps/webapp/src/pages/api/admin/mcp-tokens/{index,issue,revoke}.ts - Verify: smoke test via curl with org-admin Clerk session
-
T16 (P1, human: ~30min / CC: ~5min) —
.github/workflows/deploy-ops-mcp.yml+ update todeploy-branch-gate.yml- Surfaced by: A5/D6
- Files:
.github/workflows/deploy-ops-mcp.yml,.github/workflows/deploy-branch-gate.yml - Verify: PR into gcp-staging triggers the branch gate on changes under
apps/ops-mcp/**
-
T17 (P1, human: ~15min / CC: ~5min) — BCC CSR inbox on every order-confirmation email
- Surfaced by: D18 outside-voice (Mac visibility)
- Files:
apps/ops-mcp/src/tools/email-resolution.ts(BCC append) - Verify: test asserts BCC header present + configurable per org
-
T18 (P1, human: ~30min / CC: ~5min) — Weekly cron job DELETE on
audit_logs(rows older than 13 months)- Surfaced by: D19 (audit retention v1)
- Files: cron declaration (dagster or pg_cron)
- Verify: dry-run reports rows affected; first real run completes
-
T19 (P1, human: starts after Sprint 1 ship) — Sprint 2 Fabio’s dagster pipelines for allocation silver tables
- Surfaced by: Sprint 2 data foundation
- Files:
apps/dagster/erp_pipeline/,apps/dagster/scripts/direct_*_sync.py - Verify: silver tables populated with allocation analytics data
-
T20 (P1, human: ~3h / CC: ~1h) — Sprint 2 MCP tools (
inventory_snapshot,open_orders_for_skuvia Typesense;customer_fill_rate,recent_shorts*,record_allocation_decisionvia Postgres silver)- Surfaced by: Sprint 2 tools list, D7-revised-2
- Files:
apps/ops-mcp/src/tools/*.ts - Verify: per-tool tests pass + latency instrumentation logs to audit_logs.metadata
-
T21 (P2, human: ~30min / CC: ~10min) — Per-tool latency instrumentation to
audit_logs.metadata.latency_ms- Surfaced by: D7-revised-2 (BigQuery tripwire requirement)
- Files:
apps/ops-mcp/src/middleware/instrument.ts - Verify: a tool call adds a row with
metadata.latency_msset
Completion Summary
Section titled “Completion Summary”- Step 0: Scope Challenge — 7 existing-infra discoveries surfaced; 1 scope reduction accepted (Office.js add-in deferred under all outcomes per D1, then reframed in D17 with David’s “Claude generates Excel” rationale)
- Architecture Review — 5 issues found, all resolved (D2-D6: graph-mcp pattern, broken email-api reference, auth model, diagrams, deploy CI)
- Code Quality Review — 3 issues found, all resolved (D7-revised-2, D8, D9: data path with Typesense/Postgres split, template safety, idempotency)
- Test Review — coverage diagram produced (~50 paths, 0 existing tests on net-new service); 5 test issues resolved (D10-D14)
- Performance Review — 1 issue resolved (D15→D19 revised retention strategy)
- Outside Voice — Claude subagent run (Codex unavailable on user’s ChatGPT plan); 6 findings consolidated into 3 tensions; 2 incorporated (D17, D18), 1 revised D15→D19; 1 challenged by user with stronger rationale (D17 Claude generates Excel)
- NOT in scope — 12 items documented with rationale
- What already exists — 8 in-repo precedents credited
- TODOS.md updates — 2 items added (admin token UI, Min pipeline integration)
- Failure modes — 1 critical gap flagged with mitigation (wrong-contact email mitigated by CSR-inbox BCC)
- Parallelization — Sprint 1: 3 lanes (A+B+D), Sprint 2: 2 lanes (E+F); no conflict flags
- Implementation tasks — 21 tasks (T1-T21), all P1 except T21 (P2 instrumentation)
- Lake Score — 5/5 recommendations chose complete option (cross-org + auth + email resolution + Temporal + DLQ alert all full coverage)
What I noticed about how you think
Section titled “What I noticed about how you think”-
You rejected four pre-baked alternatives and built a fifth from first principles. I queued up A/B/C/B-prime — all variations of “panel + workflow.” You said “no frontend at all. Just the right data foundations.” That isn’t impatience or shortcut-taking; it’s reading where the substrate is going (MCP standard, AI clients as connectors) and refusing to pour engineering into a UI you’re about to obsolete. The cost: you walked away from the “forwardable dashboard screenshot” sales asset. The trade: you bought a better story for the “we’ll build our own AI agent” objection. That’s a deliberate strategic choice, not a tactical one.
-
You let the subagent reframe P2 and kept the sequencing, then changed the surface anyway. Most founders either accept the cold read whole or reject it whole. You took the part that survived translation (“customer email is the forwardable artifact”) and discarded the part that didn’t (“therefore build a panel”). Same move you made in epic-roentgen with the rule-corpus moat: extract the load-bearing constraint, drop the surface decision attached to it.
-
Then you tightened it again to “Claude in Excel.” First pass was MCP + skills + auth, no frontend. Second pass added “but Excel is where they live — meet them there.” That’s a sharper version of the same instinct: don’t build YOUR surface, augment THEIR surface. Excel-as-cockpit is the kind of architectural call that sells differently than any dashboard. The demo video isn’t “look at our app” — it’s “look at how Sam’s existing spreadsheet just got smarter.” That’s a much stickier story.
-
You corrected me twice during the eng review with sharper reads of the actual repo state. Once on Iceberg / Typesense / Postgres splits — I had Drizzle-on-Postgres as the universal answer; you knew Typesense is the typical query layer. Once on “Claude generates Excel files” as the fallback if Claude-for-Excel doesn’t yet consume external MCP — I’d queued an architectural-purity defense; you pointed out Claude can WRITE Excel files, so we never need to embed in Excel. Both corrections shrunk the design’s risk envelope and didn’t waste cycles on architectural framing — you just stated what was true and the design got better. That’s the senior-engineer move that compresses meetings.
-
You moved through six premises in one cycle. Compare to epic-roentgen where you spent multiple AskUserQuestion rounds revising P1 and P3 individually. Either the harness frame is now load-bearing enough that you trust it as the constraint set, or you’re moving faster on customer-#2 territory because you already paid the thinking cost on customer #1. Either way, the trust-the-frame signal is real — and worth verifying in 2-4 weeks by checking whether you’d ALSO reject a UI for Genfit or VMS pain, or whether CK is the canonical-AI-client shop and others aren’t.
-
You named the build vehicle in the same breath as the surface decision. “I’ll get my factory working on it” — Symphony orchestrator. You’re not separating “what to build” from “how to execute it” anymore. That’s the founder mode that lets one part-time person ship at a rate that confuses competitors. Worth protecting: the moment Symphony starts shipping things you didn’t ask for, the loop breaks.
-
You accepted a real security debt during the CEO review and named it as a debt, not waved it off. D3 on token TTL / rotation / IP allowlist — you chose defer-entirely, explicitly. That’s the founder-mode call (“speed now, harden before customer #3 or fundraising due-diligence”). It’s the right kind of debt — small surface, single mitigation (audit log + CSR BCC), clear trigger to come back. The wrong kind would have been waving it off as “we’ll be fine”; you named it.
GSTACK REVIEW REPORT
Section titled “GSTACK REVIEW REPORT”| Review | Trigger | Why | Runs | Status | Findings |
|---|---|---|---|---|---|
| CEO Review | /plan-ceo-review | Scope & strategy | 1 | CLEAR (PLAN) | HOLD_SCOPE; strategic premise affirmed (CK over Genfit AR/AP); 2 findings (token threat-model deferred, email kill-switch added) |
| Codex Review | /codex review | Independent 2nd opinion | — | — | — |
| Eng Review | /plan-eng-review | Architecture & tests (required) | 1 | CLEAR (PLAN) | 15 issues across 4 sections, all resolved; 1 critical failure mode flagged (wrong-contact email) mitigated by CSR-inbox BCC |
| Design Review | /plan-design-review | UI/UX gaps | 1 | CLEAR (PLAN) | score 3/10 → 9/10; 5 decisions made: email IA, Short Notice email type added to Sprint 1, anti-slop blacklist, mobile/a11y standards, MCP tool response interpretation hints |
| DX Review | /plan-devex-review | Developer experience gaps | — | — | — |
- CODEX: Codex unavailable on user’s ChatGPT plan (gpt-5 / gpt-5.4 / gpt-5-codex all model-restricted); fallback Claude subagent ran the outside voice during /plan-eng-review. 6 outside-voice findings consolidated into 3 tensions; 2 incorporated as design changes (D17 Excel-add-in rationale + Sprint -1 expansion, D18 CSR-inbox BCC + 48h timeout), 1 prompted a D15→D19 revision (partitioning → composite index + cron + documented v2 tripwire). One outside-voice finding challenged by user with a stronger rationale (Claude generates Excel files as artifact path).
- CROSS-MODEL: Eng review + outside voice mostly agreed on architecture and sequencing. Subagent caught Mac-visibility gap (CSR inbox BCC) and partitioning overreach (D15→D19 revised) the structured eng review didn’t surface. User caught reality-checks both reviewers missed (Iceberg/Typesense data path; Claude generates Excel files). CEO review surfaced two findings neither prior pass caught (token theft vector + email kill switch). Design review surfaced the partial-fulfillment email gap that all three prior reviews missed.
- UNRESOLVED: 0 — every AskUserQuestion across /office-hours + /plan-eng-review + /plan-ceo-review + /plan-design-review has been answered.
- VERDICT: ENG + CEO + DESIGN CLEARED — design ready to implement. Sprint -1 starts when David fires the Slack messages (Mac + Sam) and the Anthropic-docs read. Token-hardening is the only accepted security debt; trigger is customer #3 or fundraising due-diligence, whichever comes first. CK design assets (logo, brand color, signoff) needed before email visual polish lands but Sprint 1 can ship with text-only fallback while assets are gathered.