Industry

Design systems / AI tooling

Client

Self-initiated (Reportcraft)

AI design governance. A versioned design system for AI-generated reports

AI design governance, a Reportcraft experiment

Can reusable guidance make an AI-generated report worth acting on?

Reportcraft is a self-initiated experiment. It asks whether reusable design guidance, held in a versioned design.md file and stylesheet, can make an AI-generated performance report clearer, more accessible and more useful for a team deciding what to work on next. One fixed synthetic dataset ran through every attempt. A baseline agent received the data and the task with no design guidance. Six further runs received the same data plus the current guidance, which I revised between runs rather than editing any generated report. The baseline was factually accurate but hard to act on. It buried the headline metrics under a nine-point summary, relied on tables that scrolled sideways on a phone, and left readers to assemble the planning trade-off themselves. The accepted run led with the metrics, separated evidence from interpretation, adapted dense comparisons to the viewport, and passed a recorded set of arithmetic, keyboard, print and accessibility checks.

Role

Experiment design, review and direction

Duration

5 to 21 September 2026

Tools & Technology

Claude (generation) · ChatGPT (review) · Markdown · CSS · axe-core · Chrome DevTools

Output

design.md and styles.css v0.6, plus a baseline and six guided runs

Measured outcome

6

Guided runs to an accepted report

9 → 3

Summary points, baseline to final

4x

Viewport test coverage, from two widths to eight

3

Statement types kept explicit throughout the report

100%

Displayed figures traced to source data or calculations

4

Demonstrated defects corrected before the final freeze

Candidate comparison table at 960 pixels
Candidate work as labelled cards at 390 pixels

The loop

Each run followed the same sequence. Generate a fresh report from the current inputs. Freeze it, with copies of its inputs, as evidence of what that guidance version produced. Check transcription, arithmetic, responsive behaviour and accessibility. Review the hierarchy, clarity and planning usefulness. Then transfer what was learned into the next version of the guidance and generate again. The rule that made the experiment worth running was never repairing a generated report into the desired result. Every improvement had to be expressed as a reusable constraint that the next run would receive, or it did not count.

Prototyping the hardest section

By version 0.5 the report structure had stabilised and one weakness remained: the planning section listed candidate work but did not make the trade-off easy to evaluate. Rather than keep regenerating a full report to test one section, I isolated it as a prototype. That let me settle a desktop comparison table, labelled cards below 60rem, a separate capacity panel and per-discipline scenarios before folding the pattern into the guidance. Version 0.6 then applied it in a clean full run.

Separating evidence from interpretation

The rule that changed the reports most was making statement types explicit. Every claim is labelled an observation, a hypothesis or a proposed action, and the three labels appear once in a compact key rather than in a section explaining how to read the report. Hypothesis headings must themselves express uncertainty. A confident heading with a cautionary badge underneath is not enough, because the heading is what gets remembered. Shortlists stay illustrative unless a recommendation was actually requested, so the report never invents approval, ownership or certainty it does not have.

What I took from it

An agent will reliably reproduce a complex dataset. It will not automatically produce the most useful artefact for the conversation the data is meant to support. The difference came from explicit, reusable rules about hierarchy, evidence language, responsive behaviour and verification. The second lesson is about review. It is easy to accept a generated interface because it looks finished and its numbers check out. Keeping the generating agent's recorded checks separate from what a reviewer actually verified was the part that kept the result honest, and it is the habit I have carried into client work.

What this shows, and what it does not

The useful outcome is not the final report. It is the versioned reporting system and the traceable process used to develop it. The experiment does not isolate the effect of the Markdown guidance from the stylesheet, because both evolved together. Run-to-run model variability was not controlled through repeated identical trials. The dataset and the product are fictional, so this evaluates reporting and planning design rather than a real product decision. Axe results support the accessibility review but do not establish WCAG conformance, and the final run included no screen-reader test. Division of labour matters to how much weight these results carry. Claude generated each report and recorded its own browser, arithmetic and accessibility checks. ChatGPT reviewed the supplied HTML, guidance and notes, and challenged hierarchy, wording and planning usefulness. I judged the outputs, supplied visual observations, and controlled when the guidance changed. The reviewer did not independently reproduce every browser test.

Steve Hobbs ©

Steve Hobbs ©

Steve Hobbs ©