Read the Overview tab
Every heading and stat on this tab is dynamically named after your goal label — a Signal measuring "Smooth handoff" shows "Avg Smooth Handoff Rate," one measuring "Resolved by AI" shows "Avg Resolved By AI Rate," and so on. If you didn't mark a goal label, Gladly uses the first label instead.

Overall Performance — a row of four stats: the average rate for your goal label, total sessions evaluated, number of runs, and the date range covered.
"[Goal label] Over Time" (e.g. "Smooth handoff Over Time") — appears once the Signal has more than one run. A line chart with one line per label, so you can see every outcome's session count per run, over time, at once. A legend below the chart maps each color to its label. Hover a point to see that run's counts across all labels; click a point — or use the table below the chart — to drill into the sessions behind that specific outcome.
Below the chart, a table lists every run as a row, with one column per label showing how many sessions landed in that label for that run, under the caption "Select an outcome at a run to view its sessions."
[Goal label] Occurrence Rate by Agent — a stacked bar per agent (or guide), showing the full proportion of every label side by side, not just the goal label. The criterion's description appears just above the chart as a reminder of exactly what's being measured, and a legend at the bottom maps each color to its label.
Read the Run List tab
Switch to Run List to see every individual run.

Occurrence rate measured on — a dropdown letting you choose which label the Occurrence Rate and Trend columns are calculated against (defaults to your goal label, but you can pivot to any label in the criterion).
Individual Runs table — Date, Session Date Range, Sessions (count evaluated), Occurrence Rate (% of sessions matching the selected label), and Trend (the percentage-point change vs. the previous run, with an up or down arrow). A run that was interrupted shows a Cancelled tag under its date.
Click any run to drill into the sessions it evaluated.
Review individual sessions in a run
Opening a run shows every evaluated Conversation, its date, Channel, AI-assigned label, and reasoning. The header tracks Total, Reviewed, and Pending Review counts so you always know how much has been validated.

Use the Review action to approve or flag the AI's label. Approved sessions count toward the Reviewed total.
Click into any conversation to open the review panel: the full transcript alongside the AI's evaluation for each criterion — the label it assigned, its written reasoning, and the supporting quote (evidence) from the transcript.
Use Do you agree? to approve (thumbs up) or flag (thumbs down) the AI's assessment. Disagreements help refine your criteria over time.
Click the open conversation link to open the full session in AI Sessions for complete context on the individual Conversation.

Search for distinct AI labels
Start your review with sessions where the AI label surprises you. These often reveal where your criterion description needs refinement, or where the AI picked up on something unexpected.
Signals only evaluates the messages exchanged between the AI and the Customer, handoff metadata and handoff summaries are not included in what the AI reads.
Follow best practices for reviewing
0% isn't a failure. It just means no sessions in scope matched the goal label. Check the label you're measuring on and your filters if a result looks unexpected.
Expect about 90% accuracy. The remaining sessions are typically the same nuanced, ambiguous conversations that people would also disagree on — you don't need every session resolved perfectly to get a reliable signal.
Treat a Signal as iterative, not static. Adjust the criteria or filters as you learn, or spin off a new Signal once a pattern surfaces from your review — for example, moving from a general sentiment Signal to a targeted one like "Was a discount code mentioned?"
Editing a criterion never rewrites history. Each criterion is versioned, and past runs keep the results from the version that was active when they ran. Only future runs pick up your edits — nothing re-runs automatically.
