BYOK AI-Efficiency Stack — design (council-ratified)¶
Plain-language version: How WorkWingman Spends Your AI
Status: design ratified 2026-07-10 (llm-council fable tier, unanimous consensus; report:
llm-council/reports/2026-07-10_124945_design-sitting-workwingman-byok-ai-efficiency-st.md;
house gaming-PC council concurred). Build tickets: WW-66..WW-72 (TROI epic).
Parent standard: Token-ROI Metrics Standard v1.0 —
this stack is the WorkWingman adaptation. Compliance with that standard's §7 hard gate is required;
this repo's METRICS.md will be added in WW-66.
Scope rule: everything here meters and governs the user's own BYOK keys, only for WW workloads, only while WW runs. The models get fenced task prompts and nothing else; the metrics subsystem gets usage numbers and nothing else.
1. Metering (WW-66)¶
- Seam: the existing AIH provider-agnostic harness (delivered Jul 7–8,
ce6e40a) is the sole LLM door. Metering middleware (TokenMeteringMiddleware) wraps every harness call. No new gateway; the council'sILlmGatewayrole is played by the AIH harness interface. - Structural coverage guarantee (not procedural): Roslyn analyzer / CI grep fails the build on
provider hostnames (
api.anthropic.com|api.openai.com|generativelanguage|api.x.ai) or provider SDK references outside the harness namespace. Debug builds add aDelegatingHandlerasserting every LLM-host request carries a harness correlation-id. - Ledger:
ai-usage/<yyyy-MM>.events.jsonl, append-only monthly (mirrors Value Metrics event files). Sealed record — no string content field exists in the type, so resume/PII text cannot leak into telemetry:{ts, provider, model, keyIdHash, taskClass(enum), feature(enum), tokens{input,cacheCreate,cacheRead,output}, et, etcUsd, latencyMs, status, rateLimitSnapshot, routeDecision, outcomeRef(validated id)}. - ET/ETC formulas per parent standard §2. Ollama calls go through the same harness + ledger ($0 cash, tokens tracked).
2. Isolation (WW-66)¶
- Stateless API calls; PromptSafety fencing (existing) on every prompt;
tools: []hard-coded — no function-calling code path compiled into adapters. - Prompt builders consume explicit DTOs with per-taskClass field allowlists, never raw domain aggregates; redaction happens before prompt assembly.
- Responses are data, never executed; output length caps.
- Key material: harness resolves keys at call time via the existing vault key-broker. The
metrics subsystem reads ledger files only — enforced by project graph:
Metricsproject has no reference to the vault project.
3. Plan/tier awareness (WW-67)¶
- Declared: at key setup, user picks plan from dropdown (per provider: e.g. Anthropic Free/Pro/Max/API-PAYG; OpenAI Plus/Pro/API tier; Gemini free/AI-Pro). Stored in settings JSON, not the vault (tier isn't secret; vault keeps secrets only). Per key: alias, provider, declaredTier, userMonthlyWWCap, reservePercent (personal-chat headroom WW may never touch), model allowlist.
- Observed: adapters capture
anthropic-ratelimit-*/x-ratelimit-*/ 429 retry-after intoPlanObservationsnapshots; inference compares observed limits vs declared tier's expected. - Mismatch surfaces in UI ("headers suggest Pro, you declared Max — fix?") — never silently
overwritten. Every plan figure carries a confidence tier:
declared | header-confirmed | inferred | unknown. - %-of-window = ET consumed in provider window ÷ tier's modeled capacity. Capacity table = editable JSON with source URLs, labeled MODELED.
4. WW usage governor (WW-67)¶
Thresholds computed against the WW cap = userMonthlyWWCap × (1 − reservePercent) — WW
throttles itself before touching the user's personal quota. Per key, per plan window:
| Tier | Threshold | Behavior |
|---|---|---|
| GREEN | <60% WW cap | everything runs |
| YELLOW | 60–80% | pause overnight enrichment, study suggestions, council; user-initiated tailoring + core apply continue; prefer local/cheap routes |
| RED | 80–95% | user-clicked actions only, single cheapest capable model, banner |
| BLACKOUT | ≥95% or hard 429 | all WW AI paused for that key; failover to other keys only if user opted in; user override shows cost warning |
Fixed degradation order: enrichment → suggestions → council → tailoring-assist → core apply last.
5. Router (WW-69)¶
Task classes → models by fit + remaining budget, governor-aware. Defaults: parse/extract/classify/
dedupe → local Ollama else cheapest key; job-fit/ranking → cheap cloud; resume prose/interview
answers/cover letters → strongest key; study suggestions → mid. User-visible + overridable
(RoutingRulesGrid). Savings: MEASURED = cache reads + calls actually served at $0/cheap with
a priced same-taskClass same-window baseline; MODELED = "strong model would have cost X".
Never summed (parent §5).
6. WW caveman — compression boundary (WW-70)¶
Boundary = artifact type, not call site. taskClass (required enum on every harness call)
gates it: compression allowed for {parse, extract, classify, dedupe, score, summarize-internal};
artifact-producing classes (resume, coverLetter, thankYou, interviewAnswer, learningPlan)
structurally reject the compressed pipeline — user-facing artifacts are always full quality.
Savings = MODELED unless later A/B-measured.
7. Metrics tab + durable outcomes (WW-68, WW-71)¶
Coordination with Value Metrics (
docs/technical/value-metrics.md, in flight as of 2026-07-10 — seeTROI-COORDINATION-NOTE.mdat repo root): the/metricsdashboard belongs to the Value Metrics section; WW-68 ships as a section inside it, not a second page. The AI ledger is a separate shard family (ai-usage-YYYY-MM) reusing the MetricsEventLog append-only/never-throws mechanics — never writing intometrics-events-*. WW-71 consumes Value Metrics events (apply.submitted,application.status.changed, futuredocument.*adoption events) via stable ids foroutcomeRef, and reads apps-per-attention-hour from the summary endpoint rather than recomputing it.
Headline: $ per QUALIFIED tailored application submitted and not withdrawn within 7 days —
"qualified" = the Value Metrics definition (qualifiedApplicationsSubmitted: submitted AND
deterministic skill-coverage Good+, re-denominated by the 2026-07-10 Kerr sitting, commit
2414106). Raw submission counts are spray-gameable; the qualification gate is mechanical
(deterministic skill matching), not model-graded, so it passes the self-grading ban.
Model-graded quality multipliers (ATS score × count, quality-score-weighted adoption) remain
banned — a model grading its own output is the Kerr countable-criteria trap. WW-71
consumes qualifiedApplicationsSubmitted, activeMinutesPerQualifiedSubmission, and
responsesPerTenQualified from GET api/metrics/summary — never recomputed. fit.analyzed
events also carry adjacentCount (transferable skills, commit 6285670) — available to WW-71
for outcome-correlation analysis from event history; deliberately NOT part of the qualified
gate or coverage ratio (separate signal by design, per the Value Metrics session).
| Metric | Counts when | Gaming vector | Guardrail |
|---|---|---|---|
| App submitted + kept 7d (headline) | user-confirmed submit, not withdrawn | spam-apply burn-maxxing | response rate + apps-per-attention-hour (Value Metrics activation metric) |
| Artifact adopted | explicit user keep/export/send | auto-accept slop, tiny-edit farming | adoption = explicit action + min meaningful diff shown |
| $/outcome | — | cheap-model-everything (frugality-maxxing) | regenerations-per-artifact (rework count) |
| Interviews earned | recruiter screen logged | luck-noisy, lagging | context line only, never sole KPI |
| Learning task done | study item completed + retained | checklist spam | must map to target-role skill gap |
Tab shows per key: spend this window, %-of-plan (with confidence badge), $/durable-outcome, savings in separate MEASURED / MODELED panes, routing-decision log, governor tier banner. Real wired Angular per house policy — no stubs.
8. WW council (WW-72)¶
Multi-BYOK-model deliberation (independent answers → peer review → synthesis) for high-stakes artifacts only: resume finals, key interview answers, high-stakes cover letters. Learning plans and project descriptions eligible on user request. Never auto-convened; disabled at YELLOW+.
Consent modal before every run: "Council Mode (3 models). Est. cost: $0.15. ~15s. Window impact: 4% of Anthropic key." — per-seat list with one-click seat drop, max-spend cap, explicit "uses your keys now" label, [Approve] / [Single model instead]. Tracking: $/adopted-council-artifact — adoption MEASURED, uplift vs single-model MODELED.
9. Build order¶
- WW-66 metering + ledger + isolation enforcement (meter before optimize — no savings claims without a baseline window)
- WW-67 rate cards + declared tiers + header observation + governor
- WW-68 Metrics tab + Keys/Plans UI
- WW-69 router + visible overrides
- WW-70 compression policy
- WW-71 DurableOutcomeLinker
- WW-72 council + consent UX + adoption tracking