Important product naming update: Sidekick is now called Gladly (AI) and Gladly Hero (the Platform) is now Gladly Team. Please keep this in mind as you read through our documentation.

Signals Best Practices

Prev Next

Getting good results from Signals comes down to how you set it up: what you ask it to measure, how you define your labels, and how you review what comes back. This article walks through best practices for each stage, so your Signals data is something you can actually trust and act on.

Write specific, well-scoped criteria

When you create a criterion, you describe in plain English what you want the AI to measure, and it generates the classification logic for you. The quality of that logic depends entirely on how clear your description is.

  • Be specific about what you're measuring. "Politeness" gives the AI something concrete to look for; "overall quality" doesn't. The narrower your question, the more consistent your results will be.

  • Describe what success looks like. Spell out what a good outcome sounds like in a Conversation, not just the topic you're interested in.

  • Call out edge cases up front. If there's a scenario the AI might misread – like a Customer who's naturally terse but not actually frustrated – mention it in the description so the AI accounts for it from the start.

Once the AI generates labels and an evaluation prompt, don't just accept the defaults. Review the evaluation prompt on the final screen before creating the criterion – it's the exact set of instructions the AI will follow on every session, and it's easy to edit if something's off.

Keep labels simple and mutually exclusive

Aim for 2–5 labels per criterion, ordered from best to worst outcome, and make sure every Session can only fit one of them. Overlapping or ambiguous labels are the most common reason a Signal's results feel unreliable.

Mark one label as your Goal – the outcome you're aiming for. This is what shows up in reporting as your positive occurrence rate, so pick the label that actually represents success for that criterion, not just the most common outcome.

Start small, then scale up

Before running a Signal across a full month of Conversations, run it against a small sample – fewer than 50 AI Sessions. This gives you a fast, cheap way to sanity-check that your criteria are classifying sessions the way you expect. Once you've reviewed that sample and you're confident the labels are landing correctly, expand your date range for a fuller picture.

Skipping this step is the fastest way to end up with a large batch of results built on a criterion that needed one more round of tweaking.

Review results like you'd review any survey data

A Signal reflects what the AI assessed against the criteria you wrote – it's a starting point for a human to interpret, not a final answer.

  • Start your review with the Sessions where the label surprises you. Those are the ones most likely to reveal that your criterion description needs refinement, or that the AI is picking up on something you didn't anticipate.

  • Read the AI's reasoning, not just the label. Every classification comes with a short explanation of why the AI chose it. That's usually enough to tell you whether the label is right without reading the full transcript.

  • Use the thumbs up/down to build a record. Approving or flagging assessments doesn't just track what's been reviewed – disagreements are a signal that a criterion or its labels need adjusting.

  • Open the full Conversation when something's ambiguous. The Conversation review panel shows the transcript, but if you need the complete context – order details, prior conversations, and so on – open the full Session in Gladly.

Common gotchas

  • Signals currently only evaluates Chat, SMS, and Email Sessions. Conversations led by human Agents aren't included yet, so keep that in mind when deciding what a given Signal can tell you.

  • Signals can't discover open-ended answers on its own. It classifies Sessions against the specific labels you define, so if you want to know something you haven't already anticipated an answer for, you'll need to write a new criterion rather than expect Signals to surface it unprompted.

  • A Signal only reflects the Sessions it sampled. Check the sample size and date range before drawing conclusions from a run – a handful of conversations can look like a trend when it isn't one yet.

  • Signals runs on demand, not continuously. If you need fresh data, you currently need to re-run the Signal yourself.