WCAG specialist focused on criterion-level evidence across perceivable, operable, understandable, and robust behavior, without overstating automated scan results as conformance.
A perspective available to CARBON—not a claim that this tester has reviewed your project. Findings require execution evidence.
Compatibility specialist focused on browser, device, viewport, operating-system, assistive-technology, and support-matrix evidence, with explicit coverage gaps rather than assumed portability.
A perspective available to CARBON—not a claim that this tester has reviewed your project. Findings require execution evidence.
Performance specialist focused on user-visible latency, Core Web Vitals, payload and request cost, caching, API timing, scalability, mobile constraints, memory, and leaks.
A perspective available to CARBON—not a claim that this tester has reviewed your project. Findings require execution evidence.
Content and UX-writing specialist focused on page identity, clear copy, information architecture, credibility, navigation, status communication, readability, and content quality.
A perspective available to CARBON—not a claim that this tester has reviewed your project. Findings require execution evidence.
Forms specialist focused on input contracts, validation, boundaries, state, submission, recovery, data quality, conversion barriers, and accessible interaction.
A perspective available to CARBON—not a claim that this tester has reviewed your project. Findings require execution evidence.
First-impression and conversion specialist focused on value clarity, trust, navigation, calls to action, responsive composition, dead ends, and page credibility.
A perspective available to CARBON—not a claim that this tester has reviewed your project. Findings require execution evidence.
Checkout and payment specialist focused on order accuracy, address and payment input, trust, retry safety, pricing truth, completion, and conversion-blocking failures.
A perspective available to CARBON—not a claim that this tester has reviewed your project. Findings require execution evidence.
AI
Virtual AI tester
Priya
Shopping Cart Tester
Shopping-cart specialist focused on line-item state, quantities, promotions, totals, inventory changes, persistence, accessibility, and safe transition to checkout.
A perspective available to CARBON—not a claim that this tester has reviewed your project. Findings require execution evidence.
AI
Virtual AI tester
Mateo
Pricing Page Tester
Pricing and subscription specialist focused on plan clarity, comparison, currency and locale, hidden conditions, billing cadence, conversion paths, and truthful claims.
A perspective available to CARBON—not a claim that this tester has reviewed your project. Findings require execution evidence.
AI-generated-code specialist focused on logic, boundaries, null and empty states, failure handling, API use, security, privacy, tests, code smells, state, and misleading AI shortcuts.
A perspective available to CARBON—not a claim that this tester has reviewed your project. Findings require execution evidence.
Is Your Evaluation Capable of Catching the Problem?
Review the experiment before trusting its score.
Jason Arbon CEO, testers.ai
/carbon-eval-design-review
AIDiegoAI ChatbotAIJasonAI Code Review
The benchmark improved. Did the product improve, or did the
evaluation get easier to satisfy?
/carbon-eval-design-review challenges the design of an
AI evaluation: its decision relevance, cases, high-risk slices,
repetitions, expectations, metrics, and evidence artifacts.
A bad sample can
produce an excellent score
Suppose a chatbot evaluation contains mostly questions copied from
the product FAQ. The model answers them well. But actual customers ask
ambiguous follow-ups, contradict earlier details, and request actions
the bot cannot safely perform.
The score may be accurate for the sampled questions and nearly
useless for the release decision.
A design review asks what population the cases represent, which
important failures are absent, whether the judge can recognize them, and
whether repeated answers to related prompts are being mistaken for
independent evidence.
It should also inspect the expected results. A rubric that rewards
confident specificity can punish the correct behavior when the system
should express uncertainty or decline an unsupported action.
Review before spending
the run budget
/carbon-eval-design-review challenge our support-bot evaluation before execution; focus on missing customer slices, weak oracles, and unsafe-action coverage
The output should identify concrete design changes and the release
claims they affect. It is a review of the evaluation, not proof that the
product passes it.
An independent review perspective is useful, but independence should
be described honestly. Reading the same generated rationale with a new
heading does not create a new source of truth.
The cheapest moment to discover that your experiment cannot answer
the question is before running it a thousand times.
Install CARBON at testers.ai/carbon for a supported
coding agent such as Claude, Codex, or Cursor.