Files
website-evaluator/FUTURE.md
2026-04-17 23:52:40 +12:00

1.9 KiB

Future Ideas — website-evaluator

M2: Claude API Analysis Layer

Problem: The report is structured data — Claude still has to read it and form recommendations manually each time.

Idea: Add an optional --analyze flag. After generating the report, call the Claude API with the report + screenshots (vision) and append an AI-generated "Recommendations" section with prioritised, actionable fixes.

Open questions:

  • Should the recommendations overwrite a section or be a separate file?
  • Which Claude model for analysis? Sonnet for cost, Opus for depth?
  • How to handle the screenshot as base64 input to the API?

Depends on: M1 verified and stable.


M3: Comparison Mode

Problem: It's hard to tell if improvements actually moved the needle without a before/after view.

Idea: Accept two URLs (or a URL + cached baseline) and produce a diff report highlighting changes in scores, new/resolved issues, and visual diffs between screenshots.

Open questions:

  • Store baselines as JSON in the output dir? Or accept two full report.md files?
  • Visual diff: pixel-diff the screenshots or just note score changes?

Depends on: M1 stable. M2 optional.


Batch Mode

Problem: Evaluating 10 pages of a site requires running the container 10 times.

Idea: Accept a newline-delimited list of URLs and produce one report per URL plus a summary roll-up.

Open questions:

  • Output structure: one folder per URL? Or a single multi-page report?

Depends on: M1 stable.


CI Integration Helper

Problem: Running the evaluator manually is fine for one-offs but doesn't scale to catching regressions in CI.

Idea: Produce a machine-readable scores.json alongside the Markdown report. Add a --fail-under flag that exits non-zero if any Lighthouse category drops below a threshold. This makes it usable as a CI gate.

Depends on: M1 stable.