Different perspectives. One harness.

Your virtual AI testing team

AI-generated testing perspectives, not human reviewers. Select a profile to explore its focus.

← All articles
Confidence & decisions · 2 min read

Evaluate the AI Behavior, Not the Demo

Turn a promising answer into a repeatable investigation.

/carbon-eval

Diego — virtual AI testerAIDiegoAI ChatbotJason — virtual AI testerAIJasonAI Code Review

The assistant gave a good answer when you tried it. Then you tried it again and it gave a different one.

That is not automatically a defect. It is a reason to evaluate the behavior over the population and conditions that matter.

/carbon-eval helps design, implement, run, and analyze evaluations for prompts, models, retrieval, ranking, agents, and other variable AI features.

Choose the population before the average

A retrieval assistant may answer common questions well and fail when documents conflict. A personalized recommendation can look sensible for the default profile and become inappropriate for a less common one.

Useful evaluation defines the intended population, important slices, expected behavior, unacceptable outcomes, and repetition strategy. A convenient collection of prompts is not automatically representative.

Exact checks are valuable where the contract is deterministic. Semantic judgments need a defensible rubric and some examination of the judge itself. A model confidently rating another model is not independent ground truth by default.

Keep the experiment attached to its version

/carbon-eval compare the current and proposed retrieval prompts on the same approved cases; include conflicting sources and repeated runs

The report should preserve the relevant prompt, model, data, retrieval, tool, and evaluator context. Show variation and severe failures rather than hiding them behind a single mean.

Running an evaluation is subject to access, data-sharing, and cost permissions. A generated harness is still unexecuted until it actually runs.

The question is not whether AI can produce a pleasing example. It is whether the behavior remains useful under the conditions your product will encounter.

Install CARBON at testers.ai/carbon for a supported coding agent such as Claude, Codex, or Cursor.

— Jason Arbon, CEO testers.ai