Skip to content
ProductOpen console

Testing and publishing

Three console surfaces sit between a change and a customer: Preview, Testing and Evals, and Publish. They do different jobs and it is worth knowing which one answers which question.

Surface Answers
Preview What does this feel like to talk to?
Testing and Evals Does it still get the answers right?
Publish Is it live, and what changed?

A conversation with your unpublished agent. Use it the way a customer would: type badly, change your mind halfway, ask the awkward pricing question.

Preview is for judgement, not proof. It tells you whether the tone is right and whether the agent rambles. It cannot tell you that yesterday’s fix did not break last week’s answer, because you will not remember to check.

This is where the checking becomes repeatable. You write cases, and a case has three parts.

User turns. What the customer says. One line for a simple check, several for a conversation where the agent has to hold context.

Assertions. What has to be true about the reply. This is the part that turns a vibe into a test. “Mentions the deposit amount” is an assertion. “Sounds good” is not.

Seed context. Memory the agent should already have, so you can test how it behaves for a returning customer without staging a real one.

Run the suite and each case is scored, with the reply and the reason kept so you can open a case and read what actually happened rather than guessing from a number.

The temptation is to write cases for the things you know work. Those pass forever and teach you nothing.

The cases worth having come from the inbox. Every handover where the agent should have coped is a case waiting to be written, already phrased the way a real customer phrased it. Turn the fix into a case at the same time you fix it, and that bug cannot come back quietly.

A case can be run across models, and you can open one to read what each replied.

This is the honest way to answer “should we switch models”, which is otherwise decided by whoever read a benchmark most recently. The bigger model is not automatically better at following your brief, and in live chat a slower correct answer loses to a fast good one more often than people admit.

Publishing takes the agent you have been editing and makes it the one answering customers. Until you press it, your edits are not live. That is what lets you leave a half-finished brief overnight.

Publish also keeps release history, so you can see what went live and when. Every publish of the brief stores a version, which means rolling back is restoring a version rather than reconstructing one from memory.

An agent can be unpublished as well as published. Unpublishing is the blunt instrument for “stop answering right now”, and it is the correct move when something is badly wrong. Fixing a live agent while it is still talking to people is how a small problem becomes a long afternoon.

  1. Change one thing.
  2. Preview it, to see whether it reads right.
  3. Run the suite, to see whether anything else broke.
  4. Publish.
  5. Add a case for whatever you just fixed.

Step 5 is the one everyone skips and the only one that compounds.

Omazy CX documentation. One voice. Every channel. Always on.

Founded by Mosthofa Imran · imran@omazy.ai