Metrics strategy and venture evidence¶
How WorkWingman measurement systems connect to venture evidence: predeclared metrics, thresholds, falsifiable assumptions, and what product telemetry is allowed to prove.
Primary sources: docs/business/venture-evidence-ledger.md, docs/testing/user-testing-telemetry-spec.md, docs/technical/value-metrics.md, docs/discovery/ww-117-onboarding-concierge-v0.md, METRICS.md, docs/technical/byok-efficiency.md.
1. Strategy principle¶
Measure what can change a decision; never let a dashboard invent demand.
| Layer | Job | Authorized strength |
|---|---|---|
| Product event logs (Value Metrics, WING-208, AI usage) | Observe behavior and engineering outcomes | Can support Observed (and carefully defined Outcome) only with explicit metric definition, numerator/denominator, cohort, window, and limitations |
| Interviews / debriefs | Motive, language, workflow stories | Interview — never merged with telemetry into a single “proof” blob |
| Pilots with scarce commitment | Signed intent, paid plan, retained use | Commitment / Outcome |
| Models, pitches, council opinions | Planning and clarity | Assumption until tested |
Rules from the venture-evidence ledger that govern metrics strategy:
- Falsifiable hypothesis before evidence collection (segment, behavior, threshold, time box, kill criterion).
- What people say ≠ what they do.
- Prefer the smallest ethical test that can reverse a decision.
- Protect participants; minimize collection; keep raw identifying notes out of the repo.
- Preserve local-first / user-owned-data promises — no covert tracking or selling candidate data as a test method.
- Product telemetry must not be treated as customer-discovery testimony about why someone bought.
2. App integration boundary (do not blur)¶
The Markdown ledger is the operational contract for business claims today.
- Do not overload the append-only
MetricsEventstream with interviews, hypotheses, or board decisions. - If an in-app venture-evidence tool is built later, it must be a separate aggregate/API: stable claim IDs, owners, timestamps, provenance, threshold revisions, consent, decision history.
- That system may reference immutable metric snapshots; it must not mutate telemetry or infer demand from usage events alone.
Presentation generators must keep observed metrics, external research, interpretation, and forecasts visually distinct.
3. Product metric families and business use¶
3.1 Value Metrics — owner hunt as marketing proof¶
Business intent: the owner’s own hunt becomes defensible proof of time and quality — not vanity volume.
Strategic KPIs (local, Kerr-gated):
- Qualified applications submitted (fit-gated)
- Automation active minutes (park-aware)
- Estimated time-saved range (baseline − active; low bound headline)
- Response rate and responses per 10 qualified
- Selectivity (fit-analyzed, deliberately not pursued)
- Offers as a first-class funnel stage
Export contract for presentations (value-metrics.md): any claim citing product telemetry needs immutable event types, qualified definition, numerator, denominator, cohort, window, baseline/calibration state, export timestamp, and limitations.
Not authorized alone: willingness to pay, organization fit, market prevalence.
3.2 WING-208 — user-testing usage (not value)¶
Business intent: learn how external testers behave and why (via moment surveys), without rewarding click farming.
Predeclared assumptions (must stay frozen before sessions):
| Assumption | Metric | Threshold / decision |
|---|---|---|
| Testers prefer Flow (simple) mode for real work | mode_time active split |
If ≥60% active time in Studio/power, simplify-first roadmap is wrong → revisit |
| LinkedIn is the dominant onboarding source | onboarding_source |
If Indeed+USAJobs dominate connections, reprioritize source polish |
| Recommended jobs drive live runs more than saved | live_run_start entry split |
If saved ≥ recommended, ranking isn’t earning trust → iterate ranking |
| Live runs mostly complete once started | live_run_end outcomes |
Abandon rate >40% → UX/reliability blocker before scaling testers |
| Surveys at moments get answered | answer vs skip per trigger | Skip >75% on a trigger → remove/retarget trigger; don’t nag harder |
Reporting rules: observed numbers with denominators; confounders (small N, invited testers, demo sessions) stated alongside. No invented traction.
Cloud sink (user-testing builds): BigQuery ww_telemetry + Looker Studio report (WING-208 Done on board). Operator usage-insights only.
3.3 AI usage / TROI — meter before optimize¶
Business intent: honest BYOK efficiency; later $/outcome claims only after a metered baseline.
- WW-66 ships measurement only (calls, latency, tokens when real).
- Future TROI work (rate cards, routing, durable outcomes) must label MEASURED vs MODELED costs and never sum them dishonestly.
- Durable outcome linkage (WW-71 design) aims at spend → submitted+kept / artifact-adopted — feeds Metrics tab + Value Metrics types; treat as design / partial until shipped end-to-end — TODO(verify) current production completeness of WW-68/71 surfaces.
3.4 Painted-door / demand observation in-product¶
Example already in taxonomy: rewards.compareOffers.clicked — observe demand for cross-offer comparison before building the full feature (threshold gate described in total-rewards design; do not invent a cleared threshold here) — TODO(verify) numeric go-signal status.
3.5 Onboarding concierge experiment (WW-117 / discovery)¶
Predeclared before sessions (docs/discovery/ww-117-onboarding-concierge-v0.md):
| ID | Hypothesis | PASS | FAIL consequence |
|---|---|---|---|
| A1 | Unaided user reaches fit-scored saved job in <15 min | ≥4/5 | Hub justified; reorder around stalls |
| A2 | Vault + LinkedIn is the drop-off cliff | >50% stalls on that step | Else reorder cards |
| A3 | User can restate mental model after 60s First Flight | ≥3/5 right | Rewrite narrative; re-test copy |
| A4 | Checklist hub beats forced wizard | ≥2/5 deviate and recover | If none deviate, linear wizard may win |
Decision table maps PASS/FAIL combinations to persevere / iterate / pivot / stop. N=5 is qualitative signal — no over-claiming percentages beyond predeclared thresholds.
4. Venture evidence ledger — claim IDs that metrics must serve¶
Stable IDs live in docs/business/venture-evidence-ledger.md. Metrics strategy must not upgrade their status without new evidence.
| ID | Claim (short) | Status (as of ledger 2026-07-20) | Metrics / test link |
|---|---|---|---|
| VE-0001 | Working product + engineering verification exist | Supported | Repo/tests/CI inspection — not product telemetry |
| VE-0002 | Job seekers complete a guided pilot | Proposed | Consent pilot; task telemetry + anonymized interviews; ≥10 starts, ≥6 core completions, ≥5 follow-ups in four weeks |
| VE-0003 | Proposed pricing sustains free core | Proposed | Thresholds not yet declared — no pricing claim advances until set |
| VE-0004 | Demand and traction | Proposed | No authorized traction wording; hypothesis TBD |
| VE-0005 | Adaptive Solutions pilot completion gate | Proposed | ≥3 of 5 complete one real qualified application; no stop-condition incident; ≥3 interviews with corroborating behavior |
Evidence strength ladder for metrics-backed claims: Assumption → Interview → Observed → Commitment → Outcome.
5. Kerr lens (strategy, not just UX copy)¶
Reward systems fail when they pay for A (easy counts) while hoping for B (good jobs via quality applications; honest learning).
Standing strategic guardrails:
- Never rank features by raw click count without pairing task outcomes.
- Never target survey completion rate.
- Never treat study link clicks as study hours.
- Never treat automation volume without qualified gating as marketing proof.
- Never treat MODELED AI savings as MEASURED cash.
- Never present invited-tester usage as general market traction.
6. Operating cadence (recommended from existing docs)¶
| Cadence | Action |
|---|---|
| Before any tester cohort | Freeze WING-208 assumption table + survey triggers; enable UserTesting only on intended builds |
| During pilot (VE-0002 / VE-0005) | Track predeclared numerators only; log confounders; keep interview notes out of git |
| Before investor/partner deck regen | Trace every demand/pricing/outcome sentence to a ledger ID; drop expired/review-due claims |
| After AI metering baseline window | Only then introduce cost cards / router savings with MEASURED vs MODELED labels |
| On contradiction | Record Mixed/Refuted; change roadmap — do not relabel the metric |
7. Open items (explicit)¶
- In-app
VentureEvidenceaggregate: specified boundary only, not implemented — TODO(verify) if a ticket exists beyond the ledger note. - VE-0003 / VE-0004 thresholds: undeclared — pricing and traction claims remain presentation-prohibited beyond hypothesis language.
- Public “no telemetry” pitch line vs WING-208 user-testing mirror: unreconciled — TODO(verify) legal/marketing review.
- Crash-report opt-in: pitch-only — TODO(verify).
- Cloud log retention and dashboard access policy: TODO(verify).
- Whether Adaptive Solutions pilot has enrolled any of the 5 seats: ledger says no enrollment recorded as of evidence date — do not invent progress.
8. One-sentence strategy¶
Use local Value Metrics to prove quality job-hunt work, use gated user-testing telemetry to falsify product assumptions with predeclared thresholds, use content-free AI metering before any dollar claims, and keep venture evidence in a ledger that telemetry can support but never silently rewrite.