WCAG specialist focused on criterion-level evidence across perceivable, operable, understandable, and robust behavior, without overstating automated scan results as conformance.
A perspective available to CARBON—not a claim that this tester has reviewed your project. Findings require execution evidence.
Compatibility specialist focused on browser, device, viewport, operating-system, assistive-technology, and support-matrix evidence, with explicit coverage gaps rather than assumed portability.
A perspective available to CARBON—not a claim that this tester has reviewed your project. Findings require execution evidence.
Performance specialist focused on user-visible latency, Core Web Vitals, payload and request cost, caching, API timing, scalability, mobile constraints, memory, and leaks.
A perspective available to CARBON—not a claim that this tester has reviewed your project. Findings require execution evidence.
Content and UX-writing specialist focused on page identity, clear copy, information architecture, credibility, navigation, status communication, readability, and content quality.
A perspective available to CARBON—not a claim that this tester has reviewed your project. Findings require execution evidence.
Forms specialist focused on input contracts, validation, boundaries, state, submission, recovery, data quality, conversion barriers, and accessible interaction.
A perspective available to CARBON—not a claim that this tester has reviewed your project. Findings require execution evidence.
First-impression and conversion specialist focused on value clarity, trust, navigation, calls to action, responsive composition, dead ends, and page credibility.
A perspective available to CARBON—not a claim that this tester has reviewed your project. Findings require execution evidence.
Checkout and payment specialist focused on order accuracy, address and payment input, trust, retry safety, pricing truth, completion, and conversion-blocking failures.
A perspective available to CARBON—not a claim that this tester has reviewed your project. Findings require execution evidence.
AI
Virtual AI tester
Priya
Shopping Cart Tester
Shopping-cart specialist focused on line-item state, quantities, promotions, totals, inventory changes, persistence, accessibility, and safe transition to checkout.
A perspective available to CARBON—not a claim that this tester has reviewed your project. Findings require execution evidence.
AI
Virtual AI tester
Mateo
Pricing Page Tester
Pricing and subscription specialist focused on plan clarity, comparison, currency and locale, hidden conditions, billing cadence, conversion paths, and truthful claims.
A perspective available to CARBON—not a claim that this tester has reviewed your project. Findings require execution evidence.
AI-generated-code specialist focused on logic, boundaries, null and empty states, failure handling, API use, security, privacy, tests, code smells, state, and misleading AI shortcuts.
A perspective available to CARBON—not a claim that this tester has reviewed your project. Findings require execution evidence.
You ran a thousand prompts. That sounds more convincing than a
hundred.
Unless they are ten templates repeated a hundred times, the old and
new models saw different cases, and the score treats every answer as
independent.
/carbon-statistical-review examines the statistical
reasoning behind AI evaluation results: paired comparisons, clustered
inputs, repeated runs, ordinal ratings, calibration, intervals, and
practical significance.
Start with what was
actually sampled
If two models answered the same questions, that pairing matters. If
many questions came from the same customer conversation, that dependence
matters. If a rating scale is ordinal, a small numerical difference may
not mean what the chart suggests.
The reviewer should connect the analysis to the intended population
and decision. A result can be statistically distinguishable yet too
small to matter. A rare severe failure can matter even when the average
looks stable.
Keep the underlying
records available
/carbon-statistical-review inspect the paired model comparison and raw result table; check dependence, uncertainty, and practical significance
The report should identify unsupported assumptions and appropriate
follow-up analysis. If only aggregate numbers are available, it cannot
reconstruct the missing sampling history or invent a trustworthy
interval.
This command is not a license to decorate every dashboard with
significance claims. It is a way to find out whether the numerical
argument supports the conclusion.
AI can make analysis faster. It can also make a weak analysis look
professional very quickly. Review the design and the data before
trusting the precision of the presentation.
Install CARBON at testers.ai/carbon for a supported
coding agent such as Claude, Codex, or Cursor.