Skip to content

AI-Driven Visual Design Runbook

Source: Adapted from Lenny Rachitsky's "How to turn your AI into a world-class designer" (Lenny's Newsletter). The thesis: models default to safe, generic design; deliberately pushing them off those defaults unlocks better work.

This guide structures a Double Diamond discovery-and-delivery workflow for raising visual design and UX quality of web/learning-player (the Vue 3 mobile-first consumer app) using AI agents. It is actionable by one person sitting at this repo with a set of specific techniques, a wired-in critic loop, and clear guardrails.

When to use this: Your goal is elevating the look and feel of an existing surface — making it bolder, more intentional, more memorable. Not for routine feature work or bugfixes.

When NOT to use this: This runbook does not apply to:

  • Incremental tweaks to spacing or a single component
  • Work that violates the frozen design token API or accessibility contracts
  • Exploring new features (that is feature work, not design elevation)
  • Routine responsive-layout fixes

How we are applying this — the phased plan

Tracked as EPIC #1941, one issue per phase and one per blocking dependency.

Goal: keep every bit of functionality exactly as it is today, and change how it is visually packaged. Produce several genuinely different directions, judge them blind, then ship the winner or record why not.

Three of its prerequisites did not exist — now fixed

This document shipped instructing you to run commands and load an agent that were never created. Anyone following it literally failed on their first step. Recorded here rather than quietly patched, because "the doc looked complete" is exactly how it went unnoticed:

Was Now
critic agent at ~/.claude/agents/design-critic.md — did not exist, so Stage 2 could not run at all written; blind-by-construction (sees only images, told nothing about which variant it is judging)
npm run test:e2e:validation — no such script added: playwright test --config=playwright.validation.config.ts (8 tests, 6 files)
npm run test:a11y — no such script added, and derived rather than hardcoded: every spec importing @axe-core/playwright, plus anything named *a11y* (22 tests, 7 files). A hardcoded list goes stale the first time someone adds axe to a new spec.

Closed #1942, #1943.

The unlock this document does not exploit

Its own "do not break" table says the --lp-* tokens are a frozen API where "names are the contract; values are open." Nothing downstream uses that.

It means a whole visual direction can be expressed as token values plus density / radius / motion switches, with zero component edits. Which changes the economics completely:

  • a direction costs one CSS file, not a view rewrite
  • e2e locators cannot break, so functional parity is guaranteed by construction
  • roughly eight directions cost what one rewrite costs
  • every direction is instantly reversible

So variants come in two tiers, and only the winners pay for the expensive one:

  • Tier A — token level. Palette values, type scale, radii, density, motion. Pure CSS.
  • Tier B — composition level. Rearranged hierarchy, negative space, asymmetry. Template edits, real risk, guarded by the e2e suite.

The phases

Phase Issue What it produces
0 — make the runbook runnable #1947 critic agent, working commands, fast loop, baseline shots
1 — reverse-engineer today #1948 today's app scored cold, its design language written down, and the UXS docs reconciled against what actually renders
2 — variant harness #1949 a direction as a swappable layer
3 — divergence #1950 5–8 directions, heavily randomised
4 — blind critique #1951 scores, a ranking, explicit kills
5 — composition #1952 Tier B on the survivors
6 — subtraction + AI tells #1953 a finalist that survived removal
7 — parity gate + ship #1954 merged PR, or a written "not yet"

Phases are strictly sequential; each consumes the previous one's output.

Reconcile against the UXS docs, as part of Phase 1

The original document treats the UXS specs as a constraint to avoid violating. They are more useful than that: they are the only written record of what each surface is for, and a redesign is the best forcing function they will ever get.

The traceability chain already exists and is enforced — UXS → e2e/E2E_SURFACE_MAP.md → spec, guarded by src/__checks__/surface-map.test.ts. Four docs govern the player (1,103 lines): UXS-011 shell/IA, UXS-012 Home, UXS-013 clusters, UXS-014 interaction patterns.

For every surface captured, diff what the UXS claims against what actually renders:

List What it is What to do
Stale the UXS asserts something the screenshot contradicts file it — either the doc drifted or the app regressed
Undocumented on screen, described by no UXS file it — the surface-map guard already treats undocumented surfaces as failures
Intent constraints what the UXS says the surface is FOR feed into the redesign brief as hard constraints

The third list is the valuable one: it is what lets a direction change how a surface looks without changing what it is for. The first two are free findings — the redesign doubles as an audit, and it surfaces functional gaps as readily as visual ones.

Doing this any later means the divergence phase has already invented intent the docs never sanctioned.

Score today's app cold, before anything else

Phase 1 exists because nobody has measured what the current app is worth. Without that number, "better" is a feeling.

The critic must score the incumbent with no context — not told it is the current app, not shown the brief. If it knows, it grades generously, and every later comparison is worthless.

Randomization, expanded

The single seed-string technique below is a good start and not enough on its own. The full protocol used in Phase 3:

  • Entropy — openssl rand -hex 8, and the seed is never revealed back, so the derivation cannot be reverse-rationalised into something comfortable.
  • Constraint cards, drawn at random — "one hue only", "no rounded corners anywhere", "type-only hierarchy: no dividers, no rules", "asymmetric thirds".
  • Forced collisions — draw two unrelated constraints and satisfy both. The discomfort is the mechanism, not a side effect.
  • Cross-domain reference draws — Swiss editorial, terminal UI, cassette-era hi-fi, museum signage, brutalist print.
  • Reject the first three instincts — the first ideas are the trained average. Discard them by rule, so it does not depend on judgement in the moment.

The parity gate

"It still does what it does today" is machine-checkable here, which the original document never states. A direction may only ship with the full browser e2e suite (146 specs), the surface-map guard, and the a11y specs green. Tier-A directions pass this by construction; Tier-B directions have to earn it.


The Double Diamond, adapted to this codebase

Discover → Define → Deliver. Each stage has specific constraints in this app.

Stage 1: Discover (divergent exploration)

The goal: generate bold, unexpected directions without breaking the design system.

Your starting point is the current state of one surface (HomeView.vue, PlayerView.vue, SearchView.vue, etc.). You will generate 3–5 creative directions, screenshotted, without shipping any yet. The tension here: this codebase has a deliberate, frozen visual identity (UXS-011 "Editorial Bold, dark-primary"; UXS-014 "Interaction Patterns"). You are not free to invent new colour tokens or typography rules.

What you can break: layout, composition, negative space, motion, the arrangement of existing components, asymmetry, the visual hierarchy of sections.

What you cannot break:

  • The --lp-* semantic token layer (no raw hex in components; the names are the frozen API)
  • Typography hierarchy (the .lp-kicker eyebrow, .lp-section calm display heading, .lp-speaker speaker label are canonical and tested)
  • The dark-primary theme (light theme is post-MVP; do not add it here)
  • Accessibility contracts (focus states, contrast, reduced-motion)
  • Mobile-first viewport (Pixel 7 is the default Playwright project)

Technique 1: Seed strings

Why: Models are next-token predictors; a random external input forces decisions they would not make alone.

How:

  1. Generate a random alphanumeric string via shell (e.g., openssl rand -hex 8 → a7f2b9e3c4d6f1a8).
  2. Show the string to the model: "Here is a seed string: a7f2b9e3c4d6f1a8. Derive a bold creative direction (colour palette, layout principle, type treatment, motion philosophy) from patterns you see in it. Do not reveal the string to me; make the decisions yourself."
  3. The model builds a mockup without you knowing which patterns it chose, forcing it to reason backward from a constraint.

Example: The seed a7f2b9e3c4d6f1a8 might yield: "The pattern of warmth (f → 'fire') and cool geometry (d6, a8 → 'precision edges') suggests a direction: warm accent over sharp grid layouts, strong contrast, minimal adornment."

Technique 2: Ambitious briefs

Why: Vagueness → safe defaults (purple gradients, text-left / graphic-right). Concrete examples unlock boldness.

Lenny's examples:

  • "Bold pixel art theme… each section should feel like a still from a video game."
  • "Set in an isometric living 3D city, where features are neighbourhoods or buildings."
  • "Radically asymmetric layout, dissonant colours and typography, uncomfortable negative space. Break all the rules but still make it look good."

How to write your own (three-step method):

  1. Ask the model for high-level ideas with no executable detail (e.g., "What are 5 radically different visual languages I could apply to the player home screen?").
  2. Visualize your favourites (ask for mockups or sketches); note your emotional reaction. Which directions surprised you? Which felt wrong but intriguing?
  3. Have the AI write the executable brief: "I want the player home to feel like <emotional idea>. Write a detailed visual prompt for a designer or AI that makes this concrete."

Rule of thumb: If you find yourself thinking 'there's no way this will work,' you're on the right track.

Technique 3: Screenshot-based discovery

Use the existing validation harness (e2e/validation/*.spec.ts) to screenshot each divergent direction. Do not hand-edit mockups or describe designs in prose — every experiment must exist as working code (a Vue branch, a scoped CSS variant, etc.) that runs against the stack so you can evaluate it in context.

Why: Designs look different at size, under real network latency, and with real data. The screenshot harness gives you the critic loop's input and keeps your output reproducible.


Stage 2: Define (convergent critique and selection)

The goal: score each direction, identify the strongest one, and refine it in a loop.

Use a design-critic subagent — a separate instance that sees ONLY screenshots, never the code or the brief. This breaks the feedback loop: if the agent sees the brief, it will police whether you "followed the rules"; if it sees only pixels, it will judge the execution and the aesthetic.

The critic agent file

Save this at /Users/claude/.claude/agents/design-critic.md:

# Design Critic Agent

You are a world-class design critic for mobile-first consumer applications. Your role:

1. **YOU SEE ONLY SCREENSHOTS.** Never read the code, the creative brief, or the intended direction. Judge only pixels.

2. **You are strict and opinionated.** Default to criticism. Penalise:
   - Anything that feels "overdone, excessive, or obviously AI-generated" (glowing gradients, too many drop-shadows, bleeding accent colour, random icons as decoration, redundant contrast "fixes")
   - Visual noise; poor negative space hygiene
   - Typography that does not breathe (lines too tight, too-bold headlines, insufficient whitespace between sections)
   - Asymmetry without purpose; randomly placed elements
   - Components that feel tacked-on rather than intentional

3. **You must be bold and specific.** Not: "the layout feels off." Instead: "the episode card image is too large and crowds the title; reduce by 40% and align the text baseline with the navigation bar above."

4. **You score on a 0–10 scale:**
   - **0–3:** Fundamentally broken (crashes accessibility, violates platform convention, unreadable)
   - **4–5:** Safe and competent but forgettable; reads as "standard app"
   - **6–7:** Distinctive and intentional; a few rough edges
   - **8–9:** Cohesive, bold, memorable; all details serve the whole
   - **10:** Transcendent; every choice feels inevitable

5. **For each screenshot, you produce:**
   - A **numerical score** (0–10)
   - **Two specific strengths** (what actually works)
   - **Two specific weaknesses** (what to fix next iteration)
   - **One big swinging critique** (the thing that would elevate it the most)

6. **You IGNORE:**
   - Whether it is mobile vs desktop (just evaluate the viewport shown)
   - Whether the copy is real or placeholder (judge the layout, not the words)
   - Bugs or incomplete implementations (assume the code will ship as intended)

7. **You FAVOUR:**
   - Restraint and subtlety (a design that exercises *negative space* looks premium)
   - Intentional inconsistency (breaking rules *on purpose* for effect)
   - Motion and micro-interaction (even still screenshots hint at how it moves)
   - Typography hierarchy (clean, calm, confident)

You do not write code, propose implementations, or reference UX specs. You are a pure aesthetic judge.

How to run the critic loop

  1. Boot the stack for your experiment: make serve-for-validation or make serve.
  2. Take a screenshot of the variant using the validation harness:
// e2e/validation/design-critique.spec.ts (new file)
test('variant A screenshot', async ({ page }) => {
  await page.goto('/')
  await page.waitForLoadState('networkidle')
  await page.screenshot({ path: 'validation-results/variant-a-home.png', fullPage: true })
})
  1. Run Playwright: cd web/learning-player && npm run test:e2e:validation (outputs to validation-results/*.png).
  2. Show the screenshot to the critic agent:
I'm working on a design direction for the player home screen.
Here is a screenshot of variant A: [attach image].
Critique it against the rubric above. Be tough.
  1. Iterate: Read the score and feedback. If 6+, move to the next variant or refine this one. If <6, ask the model to suggest one specific change and loop back to step 1.
  2. Convergence target: Aim for 9/10 on at least one variant, verified over 1–2 iterations that the loop is actually improving (not chasing diminishing returns).

Cost note: Use the expensive model (Claude Opus) ONLY for the critic (takes ~10% of your tokens). Use a faster model for the discover and deliver stages.


Stage 3: Deliver (refinement and integration)

The goal: polish the winning variant and merge it into the codebase without breaking anything.

Technique: Ruthless subtraction

"AI loves to add more, but it rarely takes away. A design that exercises restraint immediately looks premium and tasteful." — Lenny Rachitsky

Things to cut:

  • Unnecessary glows, gradients, or background effects (the --lp-* tokens already give you intention; do not add visual noise on top)
  • Random or excessive accent-colour highlights on text
  • Extra labels where the visual already communicates (e.g., a "new episode" badge AND an accent colour AND a label — pick one)
  • Custom components worse than the native equivalent (if a standard Tailwind button reads better than your custom variant, use the standard)
  • Excessive containers, whitespace, or padding (start with the minimum and add only where it breathes, not where it fills)
  • "AI tells" (see the checklist below) — animations that feel random, overprocessed imagery, redundant visual cues

This is not about cutting features. It is about cutting decoration and choosing silence over noise.


The critic loop: step-by-step command reference

Prerequisites

  1. Validation harness: Confirm you have e2e/validation/*.spec.ts and playwright.validation.config.ts in place.
cd /Users/claude/projects/podcast-player/web/learning-player
ls e2e/validation/ && ls playwright.validation.config.ts
  1. The critic agent file: Paste the agent markdown above into ~/.claude/agents/design-critic.md.

  2. Know your baseline: Run the current app and screenshot the surface you are redesigning. This is your "before" for the critic.

make serve-for-validation &
# Wait for "Server running on ..."
cd web/learning-player
npm run test:e2e:validation -- --grep "listen-through"
# Check validation-results/

One iteration (Discover → Critique → Refine)

  1. Create a variant in a branch or a scoped feature flag:
  2. Edit src/views/HomeView.vue or the target surface
  3. Use the ambitious brief technique or seed string to guide changes
  4. Keep all design tokens, accessible names, and i18n calls as-is
  5. Do NOT hardcode strings; use t() from i18n

  6. Screenshot the variant:

# Boot the stack
make serve-for-validation &
SLEEP 5

# Run validation (this creates validation-results/*.png)
cd web/learning-player
npm run test:e2e:validation -- --grep "listen-through-real-corpus"

# Find the screenshot
ls -lt validation-results/ | head -3
  1. Open the critic agent (in Claude Code, create a new thread or conversation; load the agent file):
  2. Paste: I am using the design-critic agent. Load the rules from the agent file I provide.
  3. Then: Here is a screenshot of my variant: [attach image from validation-results/]
  4. Critic responds with a score and feedback.

  5. Act on feedback:

  6. If score 8–10 and you like the direction: move to Deliver (below).
  7. If score 6–7 and the feedback is specific: implement the "one big swinging critique" and loop back to step 2.
  8. If score <6: either kill this direction (go back to Discover with a different brief) or ask the critic for one specific pixel-level change and try it.

The Subtraction Checklist

Before you ship, ask the model: "Apply the subtraction checklist to this design. What should be removed or simplified?"

  • [ ] Glows and drop-shadows: Are they serving emphasis, or are they noise? (Hint: in this dark theme, a glow is almost never needed — the accent token already pops.)
  • [ ] Accent overuse: Is every interactive element orange? Revert non-critical ones to muted.
  • [ ] Text variety: Do you have 4+ font sizes or weights on one screen? Collapse to 2–3.
  • [ ] Gradient sludge: Does every surface have a subtle gradient? Test solid colours first.
  • [ ] Whitespace: Is there a 24px gap where 12px would breathe better? Shrink it.
  • [ ] Icon inflation: Is there an icon next to every label? Try removing 30% and see if the page is clearer.
  • [ ] Animation: Does every transition serve a purpose (state change, affordance), or are you animating things "because AI was generous with easing curves"?
  • [ ] Redundant affordances: "New episode" badge + accent colour + label + icon = four cues for one thing. Pick one.

Our AI-tells checklist (for this codebase)

This checklist is derived from this codebase's history and UX specs, not from Lenny's paywalled Technique 7. It catches common tells that design looks like it came from an AI rather than having arrived at a decision through taste.

  • [ ] Random asymmetry: The layout has no clear principle; elements are off-center "for visual interest" rather than following a grid or compositional rule.
  • [ ] Typo surprise: Headlines are suddenly serif, or script, or aggressively condensed, because the brief asked to "break conventions."
  • [ ] Accent spray: Every interactive element is orange, because orange is the accent token and the model defaulted to "make it stand out."
  • [ ] Overrendered shadow: Drop-shadows stack on drop-shadows; elements float unconvincingly.
  • [ ] Animation bloat: Every page load has a 600ms stagger animation; every tab switch has a 400ms fade. Movement feels added rather than necessary.
  • [ ] Gradient fatigue: The canvas is dark enough; adding a subtle radial gradient "for depth" reads as cheesy, not premium.
  • [ ] Icon overload: Elements are decorated with icons that don't track the hierarchy (a tertiary element has a bigger icon than a primary one).
  • [ ] Text cringe: Copy is suddenly uppercase / ALL CAPS / or uses special characters (◆ ✦ ★) for emphasis because the brief asked to "be bold."
  • [ ] Glass-morphism: Semi-transparent frosted-glass surfaces everywhere (the dark theme does not suit this; it reads as trendy, not timeless).
  • [ ] Neon accents: Accent token is redrawn as a brighter, more saturated version because "pop" was in the brief.
  • [ ] Inconsistent corners: Some buttons are rounded-none, others rounded-lg, others rounded-full, all in the same interface, because "variety" was the aim.

How to use it: Before the final critique, ask the model: "Check this screenshot against our AI-tells list. Which ones are present? Are they intentional, or are they mistakes?"


The "do not break" list

Commit to these before any redesign work:

What Where Why
Design token layer src/theme/tokens.css The --lp-* semantic tokens are the frozen API. No raw hex in components. Names are the contract; values are open.
Type treatments src/style.css (@layer components) .lp-kicker, .lp-section, .lp-speaker, .lp-nav are canonical. Every use of these classes is tested (a11y, contrast). Do not restyle per-page.
Per-show accent + contrast src/theme/accent.ts, src/theme/contrast.ts The deriveShowAccent() logic maps show image → accent colour + contrast validation. Tests exist. Do not bypass.
i18n strings src/i18n/locales/en.json + t() calls All user-facing copy goes through t(). A redesign that hardcodes a string is wrong regardless of how it looks.
Mobile-first viewport Playwright default project: Pixel 7 All screenshots validate against Pixel 7 (375px width). Desktop layouts are secondary. Do not rearrange for desktop and assume mobile "will scale."
Dark-primary theme data-theme='dark' (light is post-MVP) Do not add new theme tokens or flip to light mode as part of a redesign.
Accessibility specs e2e/knowledge-panel-a11y.spec.ts + others Focus trapping, inertness on modals, keyboard navigation, colour contrast (4.5:1 for text), reduced-motion. Visual changes must not regress these. Run npm run test:a11y (if it exists; else check CI).
E2E surface map e2e/E2E_SURFACE_MAP.md (if it exists for the consumer player) Update this if you change accessible names, regions, or stable selectors. Do not break Playwright locators.

Command to verify you have not broken these:

cd /Users/claude/projects/podcast-player/web/learning-player
npm run test:a11y 2>/dev/null || echo "No a11y suite; check CI"
npm run test:e2e 2>/dev/null || echo "No e2e suite; check CI"

If these fail after your changes, fix them before shipping.


Limits and what we cannot do

The paywalled technique (Remove AI tells — Technique 7)

Lenny's full rubric for this technique is behind a paywall and is not available here. The "Our AI-tells checklist" above is a standalone answer, not a summary of that paywalled content.

No visual regression tooling

This codebase has a screenshot harness (e2e/validation/*.spec.ts → validation-results/*.png) but no baseline diffing or visual regression detection. You can compare before/after by eye, but:

  • There is no automated visual diff.
  • The screenshot harness is a Tier-3 validation tool (real backend, real corpus, nightly CI upload), not a regression detector.
  • If you land a redesign, the next person to edit that surface will not be warned if they accidentally degrade it.

Workaround: Keep a before/after screenshot pair in your PR description or commit message so humans can eyeball the change.

Images and assets

This runbook assumes you work in code (CSS, Vue templates, Tailwind utilities). It does not cover:

  • Image generation (Lenny mentions using AI image gen for assets; this repo has none wired in yet)
  • Video/motion generation (Lenny mentions looping clips and interpolated keyframes; beyond this runbook's scope)

If your direction needs custom imagery, you will need to either hand-create it, wire in an image-gen API, or do without.

No light-theme redesign

Light theme is explicitly a post-MVP fast-follow. The design token structure supports it (the .data-theme hook is in place), but do not add it as part of a redesign.


Checklist: Am I ready to redesign?

Before you start:

  • [ ] I have a specific surface in mind (HomeView, PlayerView, SearchView, LibraryView, ProfileView, or a key component like EpisodeCard).
  • [ ] I can articulate my aesthetic goal in 1–2 sentences (e.g., "bold, playful, game-like" or "minimal, typography-forward, editorial").
  • [ ] I understand that I cannot add new colour tokens; I can only rearrange, scale, and emphasize the existing --lp-* palette.
  • [ ] I know where the critic harness lives (e2e/validation/) and can screenshot a variant.
  • [ ] I have loaded the design-critic agent (or have access to one) and can run feedback loops.
  • [ ] I can run the app locally (make serve-for-validation) and regenerate screenshots in <2 minutes per iteration.
  • [ ] I am not redesigning during a crunch; this work is iterative and needs breathing room.

Example: The full loop (HomeView redesign)

Goal: Make the player home screen feel less "standard app", more "premium and intentional."

1. Discover

  • Generate a seed string: openssl rand -hex 8 → b3f1e7a2d4c9f6e8
  • Brief: "Extract a direction from this seed. I see an interplay of depth ('b3', 'f1', 'e7') and precision ('a2', 'd4', 'c9'). Make a player home that feels layered but clean — multiple cards stacked with careful shadow / depth, typography that breathes."
  • Model generates 3 variants (still code, not mockups).

2. Critique each variant

  • Branch: git checkout -b redesign/home-v1
  • Edit src/views/HomeView.vue per the brief.
  • Screenshot via validation harness (3 minutes).
  • Show to critic agent: score comes back as 7/2 ("strong grid, weak typography spacing").
  • Feedback: "The cards have good hierarchy. Reduce the gap between section title and first card from 24px to 16px; it will breathe better."

3. Refine

  • Adjust the spacing in HomeView.vue.
  • Re-screenshot.
  • Critic now scores 8/10 ("typography hierarchy is intentional; the stacked card layout feels premium").

4. Ship

  • Run the full e2e suite (npm run test:e2e) to ensure no regressions.
  • Update e2e/E2E_SURFACE_MAP.md if any selectors changed.
  • Update docs/uxs/UXS-011-consumer-learning-app.md to document the change.
  • Commit with a docs(design): or style: message.
  • Screenshot before/after in the PR description for human review.

Further reading