Skip to content

WW-143 v0 — Onboarding Concierge Test (no-code experiment)

Status: predeclared, not yet run · Owner: Andrew · Written: 2026-07-23 Council mandate: fable council 2026-07-23 (report llm-council/reports/2026-07-23_183357_*) — "Smallest experiment ships NO code. This gates whether the hub gets built at all." Sequencing override (Andrew, 2026-07-23): build MVP-0 first (with per-step convenience skip + a testing bypass), run this experiment against the REAL walkthrough instead of a concierge simulation. Hypotheses, metrics, and decision rules below stand unchanged — the experiment validates the build rather than gating it. Honest cost of the override: if the test says "don't build the hub," MVP-0 is sunk cost; the decision rules still apply to v2. Method sources: Talking to Humans (Constable) — interview discipline; Testing with Humans — experiment structure; The Lean Startup — concierge MVP / validated learning.

"Your job right now isn't to sell, but rather to learn." — Giff Constable, Talking to Humans


1. What this experiment decides

Whether the First Flight + Setup Hub walkthrough (council-architected) gets built at all, and in what order its steps go. We run the proposed onboarding as a human-guided concierge session on the current app — no new code — and watch where real people stall, quit, or light up.

One assumption per round (Testing with Humans rule). This round tests the setup funnel, not pricing, not feature love.

2. Predeclared hypotheses & pass/fail (locked before session 1 — no goalpost moves)

# Hypothesis (falsifiable) Metric PASS FAIL → consequence
A1 An unaided new user reaches a fit-scored saved job in /queue in <15 min stopwatch, install-to-first-scored-job ≥4/5 testers <4/5 → hub build is justified; step order gets rebuilt around observed stalls
A2 Vault + LinkedIn setup is the drop-off cliff count of stalls/asks-for-help per step >50% of total stalls land on the vault/LinkedIn step if stalls land elsewhere (e.g. intake interview) → reorder hub cards before building
A3 After a 60-second verbal "First Flight" (the mental model: LinkedIn SAVED jobs → scrape → fit → tailor → you apply), user can restate the model at debrief unprompted restatement scored right/partial/wrong ≥3/5 restate correctly <3/5 → First Flight screen 3 needs a different narrative, test copy variants before building
A4 Users don't need a forced linear wizard — a checklist they can leave and return to suffices observed: does the user try to leave the flow / do something else mid-setup? ≥2/5 deviate and recover on their own 0 deviate → linear wizard acceptable, hub complexity maybe unnecessary (inverts a council premise — flag it)

Secondary observations (tracked, not pass/fail): time per step; which step provokes the first "why do I need this?"; whether the KeePass vault concept lands or scares; reaction to "never clicks Submit" promise.

3. Who we recruit (Talking to Humans rules)

  • N = 5 target (2 minimum viable). Sean + Shereeba cohort qualifies only if they haven't watched the app being built — one degree of separation preferred; if they're too close, each must refer 1–2 people from their own circle who are actively job-hunting ("fish where the fish are": people who currently have a LinkedIn account and ≥1 saved job).
  • Segment: active job seekers, Windows PC at home, LinkedIn account, non-developers preferred (developers forgive setup pain the market won't).
  • Recruit by asking for advice, not offering a demo: "I'm trying to learn how people set up job-search tools — can I watch you try one for 30 minutes?"
  • Each session ends with a referral ask (aim: leave with 2–3 new candidates).

4. Session protocol (~45 min per user, 1–2 weeks total, fixed end date)

Setup (before user arrives): clean Windows profile or reset app data dir; screen-record locally (OBS or Xbox Game Bar — file stays on the PC, tell the user it exists and that it never leaves the machine); stopwatch app; printed observer sheet (§6). Two people if possible: one facilitates, one takes notes ("assign a skilled notetaker").

Part 1 — Pre-task interview (10 min, BEFORE seeing the app — behavioral Qs first, prototype after, to reduce bias): - "Walk me through the last time you applied for a job. What did you actually do, start to finish?" (story, past behavior — never "would you") - "What's the most annoying part? Tell me about the last time it happened." - "What tools do you use today? What do you pay for, if anything?" (current spend, not willingness-to-pay) - Do NOT pitch WorkWingman. Do NOT explain features.

Part 2 — First Flight, concierge version (2 min): - Facilitator delivers the mental model verbally, once, ≤60 seconds, from THIS script (same words every session — repeatability):

"This app works off jobs you've already SAVED on LinkedIn. It reads your saved list, scores how well each job fits you, and tailors your documents for the ones you pick. It never applies for you — you always click Submit yourself. To do that it needs three things set up once: your profile, a small password vault that lives only on this PC, and its connection to your LinkedIn." - Ask mode preference naturally: "Do you want the simple view or the everything view?" Record the answer. Set it for them (concierge = we click, per their instruction).

Part 3 — Setup run (25 min cap, stopwatch running): - User drives. Facilitator = silent concierge: answers ONLY direct questions, never volunteers, never points at the screen. Every question the user asks = one logged stall (verbatim). - Task card handed to user, nothing more: "Get this app to show you your LinkedIn saved jobs with fit scores. Everything you need is in the app." - Steps they'll traverse (our map, not shown to them): intake interview (/onboarding) → connections + vault (/connections) → environment (/doctor, only if red) → scrape → /queue. - Hard stop at 25 min. Unfinished = data, not failure of the session.

Part 4 — Debrief (10 min, Talking-to-Humans style): - "Tell me what this app does, in your own words." (A3 — score restatement) - "What almost made you quit?" / "Where did you feel lost?" - "What did the vault feel like — why do you think it's there?" (trust probe) - "If you got home and had 10 minutes, what would you do first with this?" - No feature tour, no selling, no fixing their complaints live. "Interview to learn, not to validate."

5. Capture & scoring

  • Per-session log (fill immediately after): name, date/time, format, facilitator/notetaker, tester profile 1-liner, recording path.
  • Shared scorecard (one spreadsheet row per tester): per-step time; stall count + verbatim stall questions; step where 25-min cap hit (if it did); A1 pass?; A3 restatement right/partial/wrong; A4 deviation observed?; standout quote.
  • After all sessions — dump & sort: sticky-note every observation, cluster into patterns, map each pattern to a hypothesis. Patterns + judgment, not statistics — N=5 is qualitative signal, no % claims beyond the predeclared thresholds.
  • Skepticism rule: discount smiles and compliments; count only behavior ("Question overly positive feedback; expect false positives"). Positive words + failed task = FAIL.

6. Observer sheet (print per session)

Tester #___   Date ______   Facilitator ______   Notetaker ______
Mode picked:  simple / power / didn't care
STEP TIMES    intake ____  vault ____  linkedin ____  doctor ____  scrape ____  TOTAL ____
STALLS (verbatim question + step):
 1.
 2.
 3.
CAP HIT? y/n at step ______
A3 restatement: RIGHT / PARTIAL / WRONG — their words:
A4 deviation: y/n — what they did:
First "why do I need this?" moment:
Vault reaction:
Standout quote:

7. Decision rules (predeclared — persevere / iterate / pivot / stop)

Outcome Decision
A1 PASS + A2 as predicted Persevere: build MVP-0 exactly as council specced (First Flight + hub, vault card prominent with why-copy)
A1 FAIL, stalls on vault/LinkedIn (A2 confirmed) Persevere with reorder test: build hub, but run the council's named pivot — vault-behind-intent (defer vault until first scrape attempt) as the FIRST variant
A1 FAIL, stalls on intake interview Iterate before build: cut Profile card to resume-paste-only path; intake's remaining steps move post-activation
A3 FAIL Iterate copy: rewrite First Flight narrative, re-test verbally with 3 fresh users (still no code)
A4 inverted (nobody deviates, all want hand-holding) Flag to council: linear wizard may beat hub — cheaper build; re-decide before MVP-0
Testers can't even install / env breaks >2 of 5 Stop-and-fix: onboarding epic pauses; installer/doctor hardening becomes the ticket
Everything passes easily (<10 min, no stalls) Stop (don't build the hub): current surfaces suffice; spend the effort on post-activation teaching instead

8. Ethics & privacy (product promises apply to research too)

  • Recordings + notes stay on the test PC; tell the tester before recording; delete on request.
  • Testers use a throwaway or their own LinkedIn by their choice — never store their real password anywhere except the local vault they create themselves, and wipe app data after the session in front of them.
  • No invented urgency/scarcity in recruiting; we're asking to learn, and we say so.

9. What this unblocks

PASS-shaped results → MVP-0 fleet cards (council tiering): ceiling = contracts/state/metrics predeclaration; Sonnet = hub + orchestrator + JSONL event log + privacy panel; Haiku = First Flight screens/copy/deep-links. FAIL-shaped results → the decision-rule row above, not a bigger build.


Method quotes: "An interview guide is not a script." / "You might show people mockups… reactions still need to be taken with skepticism." — Giff Constable, Talking to Humans.