Skip to content

AI Generation Strategy (technical)

Plain-language version: ../plain/ai-generation-strategy.md · Business economics: ../business/ai-generation-economics.md · Diagrams: ../diagrams/ai-generation-lanes.md

Status. Council-approved v2 (fable tier, 5 seats, Cedric chair), corrected 2026-07-31. This doc is the source of truth for what WorkWingman generates with AI, on which provider, at what unit cost, and under which legal cover. It supersedes the provider-selection parts of ../ai-audio-providers.md (that doc remains as research reference).

Standing constraints (apply to every lane): indemnity-first for cloud generation; BYOK on desktop; no commissioned assets; factual content never comes from an image model; provenance metadata on every generated artifact (WING-298); on-the-rails — human approval before any submission; EMDR excluded.

Corrected unit rates (council-verified)

Unit Rate
Gemini TTS $10/M (2.5 Flash) – $20/M (3.1 Flash) output, 25 tok/s ⇒ $0.90–1.80 per audio-hour
Lyria 3 Clip (≤30 s) $0.04; Pro full song (≤3 min) $0.08 — indemnity-list status UNVERIFIED, songs-lane blocker
Nano Banana 2 (1K, interactive) $0.067; batch $0.034 (queued jobs only — a waiting user can't ride batch)
Claude Sonnet 5 $2/$10 promo → $3/$15 after Aug 31, 2026 (all math below uses post-promo). Cache write 1.25×, read 0.10×
Gemini 3.6 Flash $1.50/$7.50
Project chat turn (50k ctx, realistic caching) ~$0.02–0.03
STT (interview practice) ~$0.016/min ≈ $0.96/hr

Lanes (1–13)

  1. Flashcard diagrams — procedural spec→render, $0 marginal. No image model on factual content.
  2. Characters/mascots — Nano Banana 2, generate-once cast ≈ 200 imgs ≈ $13 (batch ~$7) one-time. AI images are likely uncopyrightable ⇒ trademark the characters for ownership. Negative-prompt guards against known IP/likenesses (indemnity can be voided by IP-steering prompts). grok CLI output = internal docs/tutorials only — never shipped product assets (consumer-sub ToS).
  3. Study songs — Lyria Pro bed $0.08 + TTS overlay $0.05–0.09 ⇒ ~$0.13–0.17/song + script tokens. BLOCKED until Lyria's presence on Google's indemnified-services list is verified (not there today as far as council could tell; preview SKUs carry weaker terms).
  4. Ambient/white noise — council overturned Lyria here: it's a music model, wrong tool for thunderstorms. Use CC0 field recordings (per-file audited, per the Freesound audit) for environmental beds — $0 and better; Lyria only for musical stems. Storage/CDN egress is not free: ~$0.12/GB — priced into infra, not "zero."
  5. Narration — TTS $0.90–1.80/hr; 20-min pack $0.30–0.60. Captions/transcripts ship with every audio artifact (nearly free — everything starts from a script; accessibility requirement).
  6. Audio overviews ("Study Cast" — never "NotebookLM", trademark) — own two-host pipeline: Gemini script + multi-speaker TTS ⇒ ~$0.25–0.45 per 10-min episode. Gemini Enterprise has a preview Audio Overview API, but build-own wins on control/cost/branding. Desktop: click out to the user's own NotebookLM, $0.
  7. Project workspaces (Claude-API based) — turns ~$0.02–0.03 post-promo with honest caching. Indemnity resolution: see decision log below (Claude PRIMARY as documented exception).
  8. Core inference / PageAgent — needs its own per-token unit model instead of a $1.50–3.00 black box; ladder unchanged (Sonnet ↔ Gemini Flash ↔ Haiku, copy-paste terminal fallback).
  9. Video — not offered v1. ($0.05/s is Veo Lite 720p only; higher tiers cost multiples.)
  10. Resume/cover-letter/JD-match/outreach/STAR generation — the actual core of the product. ~$0.03–0.08 per artifact (Gemini Flash) / ~$0.10–0.25 (Sonnet); moderate user ~30 artifacts ⇒ $1–4/mo. Moderation + PII scrubbing on all uploads.
  11. Mock interview voice loop (STT↔LLM↔TTS) — ~$1.50–2.50 per conversation-hour, own lane + own cap.
  12. Grounded research (Google Search grounding) — required for live job/company research. Published SKU disputed in council ($14 vs $35 per 1k grounded queries) — verify before pricing; ~$0.01–0.035/grounded turn either way.
  13. Ingestion/embeddings — PDF/DOCX/OCR + embeddings + vector store for resumes/syllabi/ textbooks. Cheap per unit; big-PDF ingestion tokens are not — meter it.

Per-user/month (cloud, generation only, post-promo rates)

Lane LIGHT MODERATE HEAVY (uncapped) HEAVY (capped*)
Diagrams $0 $0 $0 $0
Songs ($0.15) 1 → $0.15 4 → $0.60 20 → $3.00 10 → $1.50
Study Cast ($0.35) 1 → $0.35 4 → $1.40 20 → $7.00 10 → $3.50
Narration ($1.35/hr mid) 15m → $0.34 1h → $1.35 5h → $6.75 3h → $4.05
Ambient (CC0+egress) ~$0 ~$0.05 ~$0.25 ~$0.25
Images ($0.067) 0 5 → $0.34 25 → $1.68 25 → $1.68
Resume/docs lane 5 → $0.30 30 → $2.00 100 → $6.00 60 → $3.60
Project chat ($0.025/turn) 40 → $1.00 150 → $3.75 600 → $15.00 300 → $7.50
Mock interviews ($2/hr) 0 1h → $2.00 5h → $10.00 3h → $6.00
Grounded research $0.10 $0.50 $2.00 $1.00
PageAgent $0.50 $2.25 $6.00 $4.00
Generation subtotal ≈ $2.75 ≈ $14.25 ≈ $57.70 ≈ $33.10
+ sticky IP $2–5 $4.75–7.75 $16.25–19.25 — $35–38

*Caps (WING-299): 10 songs, 10 episodes, 3 h narration, 300 chat turns, 3 h interviews — chat + interviews + PageAgent dominate heavy spend, so caps must cover them, not just audio (council correction). Moderate roughly doubled vs draft v1 because the newly-modeled core lanes (resume, interviews, grounding) are real costs the draft ignored — the audio lanes were never the expensive part.

Total monthly at scale

Mix 50/40/10 light/moderate/heavy-capped; blended generation ≈ $10.45; +IP mid $3.50 ⇒ ~$13.95/user.

Users Variable Fixed infra (incl. egress) Total/mo Per-user
100 ~$1.4k ~$0.7k ~$2.1k $21
500 ~$7.0k ~$1.1k ~$8.1k $16
1,000 ~$14.0k ~$1.5k ~$15.5k $15.50
5,000 ~$70k ~$4.5k ~$74.5k $14.90
10,000 ~$140k ~$8k ~$148k $14.80

Sensitivity: all-Gemini inference for lane 7 cuts blended generation ~25–30% ⇒ 10k users ≈ $110–120k/mo. Fair-use caps and per-lane meters (visible, never silent) are what keep the sub margin-positive.

Caps (WING-299) and top-up packs (WING-300)

  • Caps are per-lane, per-month fair-use limits on the base subscription (the "capped" column above). Meters are always visible to the user, never silent — throttling users without telling them is off the table.
  • Top-up packs (WING-300, approved): when a lane's cap is hit, the user can buy more of that lane instead of upgrading tiers. Pack pricing targets ~50% margin over marginal cost (see ../business/ai-generation-economics.md).
  • Failed generations never consume quota. A generation that errors or is rejected by moderation does not decrement the user's cap or a purchased pack.

Cost controls

  • Caching realities (Claude lane): cache write is 1.25× input, read 0.10× — but every turn re-reads the whole cached context, and the default cache TTL is ~5 minutes, so sporadic study sessions re-write the cache often. Model chat cost with these honest assumptions (~$0.02–0.03/turn), not idealized cache hits.
  • Batch API: batch discounts (e.g. Nano Banana $0.034 vs $0.067) apply to queued jobs only. A user waiting on-screen cannot ride batch. Route generate-ahead work (character cast, pre-built packs) through batch; never promise batch economics for interactive flows.
  • Fixed-asset-ization: anything generated once and reused for everyone (mascot cast, stock ambient beds, template diagrams) is a one-time fixed cost, not a per-user cost. Prefer moving spend from per-user lanes into fixed assets wherever quality allows.
  • Per-lane meters feed both the caps UI and the unit-economics model — one metering pipeline, two consumers.

Decision log

Date Decision
2026-07-31 Claude is PRIMARY for cloud text (lane 7 + text lanes) as a documented exception to the indemnity-first rule: Anthropic's commercial-terms indemnity applies — verify at onboarding; reputational backstop noted. Gemini-primary is held in reserve as a ~25–30% cost lever (see sensitivity above).
2026-07-31 Indemnity-first governs the media lanes (images, music, TTS): Google paid APIs only. Lyria's indemnified-list status is UNVERIFIED = songs-lane blocker (lane 3).
2026-07-31 Suno BYOK is dead — no official API; wiring user keys would facilitate ToS breach. Desktop music BYOK candidates: ElevenLabs Music / Stability Audio (or drop desktop music).
2026-07-31 Ambient = CC0 field recordings, not Lyria (council overturn). Lyria is for musical stems only.
2026-07-31 grok CLI output never ships in product — internal docs/tutorials only (consumer-sub ToS).
2026-07-31 Trademark the character cast — AI images are likely uncopyrightable, so trademark carries ownership. Negative-prompt guards against known IP/likenesses stay mandatory (IP-steering prompts can void indemnity).
2026-07-31 NotebookLM / Claude Projects cannot be resold. Desktop links out to the user's own tools; cloud builds equivalents. Product name is "Study Cast", never "NotebookLM".
2026-07-31 Captions/transcripts ship with all audio (accessibility; nearly free from the script).
2026-07-31 Provenance metadata on every lane including stem-derived mixes (WING-298).
2026-07-31 Failed generations never consume quota.
2026-07-31 COPPA/minors gating is an open item — study features attract under-13s; age gating + data-handling needs a policy pass before study lanes go wide.
2026-07-31 Caps + top-up packs approved (WING-299 / WING-300).

Standing rules written into strategy (previously practice only)

Human approval before any submission; no fabricated credentials; no hidden auto-apply; no training our own models on provider output (model-extraction clauses); provenance stamps on every lane including stem-derived mixes (WING-298); TTS redistribution rights verified before enabling downloads; multi-language = per-language TTS multiplier, priced when scoped.