Improvements
Let a reviewer read your agent's conversations, suggest changes as cards, test each change against past conversations, and approve it into a new revision
Improvements is a tab on the agent page. A reviewer, running on a model you choose, reads the agent's finished conversations and asks what you'd ask: did the agent have the tools it needed, did its instructions lead it right, did it know the answer? Each problem it finds becomes a card that proposes a change: a tool to add, a line for the instructions, or advice for a person. You can try the change, test it against past conversations, and then approve or dismiss it. Nothing about the agent changes until an admin approves a card.
The Inline Judge scores single replies as they run, and Evaluations score a revision against a dataset. Improvements reviews whole production conversations after they end and says what to change.
Who can use it
Only Admins of the Application that owns the agent can see or use Improvements. That includes organization Admins and Owners. Nobody else sees the tab or the review button in Conversations. Some agents are limited:
- An external agent can't be reviewed, because it has no revisions to change.
- A managed agent (for example, one a knowledge source built) shows its cards read-only. Its changes are made wherever it's managed.
- An approval or test edits the agent's one agent step. If the agent has more than one, you'll get a message telling you to make the change in the editor.
Turn on reviews
Open Improvements → Settings (or Turn on reviews the first time). The Review settings dialog stores these settings. You can also read them with GET /v1/tenants/{tenant}/agents/{agent}/reviewSettings and set them with PUT on the same path:
| Field | In the dialog | Default | Meaning |
|---|---|---|---|
enabled | Review this agent's conversations | off | Turning it on requires a model, and records the time in enabledAt. Conversations that ended before then are never reviewed in the background. Turning reviews off and on again moves that time forward. |
provider | Reviewer model | none | One of the workspace's own models. It's checked when you save, so a model the workspace can't call is refused with the reason. |
dailyBudgetUsd | Daily budget | 1 | The most that reviewing this agent can spend per UTC day. Must be above $0 and no more than $100. |
cleanSamplePercent | — | 10 | The percentage of conversations with no warning signs that get reviewed anyway. API only. |
PUT replaces the whole settings object, so any field you leave out returns to its default. The response also includes spentTodayUsd, which the tab header shows as Today $x of $y.
What a review costs. A review makes one call to the reviewer model. It makes a few smaller calls when it needs to pick a tool, check whether a knowledge-base article already answers a question, or read what a teammate told the customer. The workspace is billed for these like any other model call. They count toward the daily review budget and never toward the agent's own budget, so reviewing can't turn customers away. If the model accepts a thinking level and you leave it unset or off, reviews run with high thinking.
Which conversations get reviewed
A conversation is the agent's runs in one session, up to a gap of four quiet hours. A run with no session, such as an automation or API call, is reviewed on its own. Every five minutes, a background review checks each agent that has reviews turned on:
- It considers only production runs on a saved revision that ended after reviews were turned on. A conversation counts as over after four quiet hours.
- Conversations with warning signs go first, ordered by how bad the signs are. Warning signs include a failed run, a quality check that failed, scored low, or sent a reply back for a rewrite, an undelivered reply, tool errors, a call to a tool the agent doesn't have, an escalation, a takeover, and a reopened ticket.
- A conversation with no warning signs is reviewed only if it's drawn in the clean sample (
cleanSamplePercent). - Each sweep reviews at most 10 conversations per agent, and checks the budget before every call. A conversation still waiting after 48 hours is skipped. A review that fails three times is given up.
- An agent that's over its own enforced budget isn't reviewed.
The background review reads each conversation once. If the conversation continues after it's reviewed, it isn't read again unless you ask.
Review a conversation yourself
In the agent's Conversations tab, the contact header shows how the open conversation was reviewed, plus a button to review it now. Both appear only when reviews are on. The review runs straight away, takes about as long as one reply, and is charged to the review budget. If it raises a card, the Improvements tab opens on that card. Cards from a review you asked for show up straight away, without waiting for more evidence.
| Status line | Button |
|---|---|
| Reviewed time ago in the background (or by you), then what it raised, with an Open card link | Review again |
| Not reviewed since the conversation went on. | Review this conversation |
| Not reviewed: nothing flagged it, and the background review reads only a 10% sample of unflagged conversations. | Review this conversation |
| Not reviewed: the daily review budget ran out before the background review reached it. | Review this conversation |
| Not reviewed: the background review reads production conversations only. (or it ended before reviews were on) | Review this conversation |
| Still going. The background review reads it after 4 quiet hours. · Queued for the background review. | Review now |
| The last review of it failed. | Try again |
The API equivalents are GET …/conversationReview and POST …:review. Name the conversation with executionRecordId (any run in it) or topicId (its latest conversation). A manual review is refused if reviews are off or the day's budget is spent.
What a review finds
Each review records findings, and a finding's category decides where it goes:
| Category (UI label) | Where it goes |
|---|---|
| Missing capability | A card that adds tools from the Fruxon catalog or your workspace, or advice if no tool can do it |
| Instructions, Tool use | A card that adds or rewrites an instruction, or advice |
| Knowledge | The agent's knowledge base, as a draft held for review. It becomes an advice card instead when a published article already answers the question (the agent didn't use it, or couldn't find it), or when the agent has no knowledge base. A conversation the base's own conversation source will read is left to that source. |
| Integration, Platform | Recorded only. A broken connection or a fault in Fruxon is never shown as a change to make to your agent. |
A finding that matches an existing card is added to that card as more evidence. The tab shows how many questions went to the knowledge base as drafts and links to the base's review queue. Nothing there is published until a person approves it. To list findings by where they went, call GET …/reviewFindings?disposition=. The values are MERGED (on a card), SENT_TO_KNOWLEDGE_BASE, KNOWLEDGE_SOURCE_WILL_EXTRACT, INTERNAL, and REJECTED_INVALID.
Cards
A card shows a title, its impact (High impact stands out, Medium impact and Low impact don't), its category, and when it was last seen. Below that come a summary and Seen in N conversations. Expand that line to list each conversation with Open conversation and, where the change can be tried, Try the change here. Then comes The change (for advice, What to do), and finally the test results. The change is one of four kinds:
| Kind | The change | Notes |
|---|---|---|
ADD_TOOLS | Attaches catalog tools, sometimes with a line telling the agent when to use them | A tool that changes things is attached with Asks a person each time: a person must approve each call. A tool whose integration isn't connected shows a Connect button instead of Approve. |
ADD_INSTRUCTION | Appends text to the end of the system prompt | Show where it goes in the instructions shows it in context. |
REPLACE_INSTRUCTION | Replaces one exact passage of the system prompt | If the passage disappears or appears more than once, the card closes. |
ADVICE | Written guidance for a person, such as turning on memory, escalation, a knowledge base, or the quality check | Can't be approved or tested. Do what it says, then click Mark done. A deploy that turns the named feature on closes the card. |
When a card appears. A new card waits as a CANDIDATE. The tab doesn't show candidates. It counts them as Watching N more patterns. A candidate becomes a card when either of these is true:
- A second conversation shows the same problem.
- One conversation on the agent as it is now shows it with high impact, and the change isn't a rewrite of a passage of instructions or a tool that changes things. Those two kinds always need two conversations.
A card from a review you asked for skips the wait. If a later revision has changed the agent since the card's conversations, the card shows May be fixed by vN.
Card lifecycle
The Waiting, Applied, and Dismissed filters group these statuses:
| Status | Filter | What moves it |
|---|---|---|
CANDIDATE | hidden | Becomes OPEN once enough evidence arrives (see above) |
OPEN | Waiting | Approve → APPLIED. Dismiss → DISMISSED. A deploy that already has the change → OBSOLETE. If a later revision changed the agent under it, and at least 14 days and 10 reviewed conversations since then don't show the problem → QUIET. |
APPLYING | — | Brief, while an approval runs |
APPLIED | Applied | Fourteen days and 10 reviewed conversations on the new revision without the problem → VERIFIED. The problem shows up on that revision → RECURRING. A deploy of an older revision, including Undo → OPEN. |
RECURRING | Waiting | Shows "Still happening after vN". You can approve again (refused if the change is still live) or dismiss it. |
VERIFIED | Applied | Same as APPLIED, if the problem comes back or the revision is rolled back |
DISMISSED | Dismissed | Five more conversations showing the problem → OPEN, unless it was dismissed as Wrong |
OBSOLETE | Dismissed | Final. Shows "Closed — reason": the agent already has the change, can't take it, or every conversation behind the card is past retention. |
QUIET | Dismissed | Comes back if the problem shows up on the agent as it is now |
Try a change on one conversation
Try it (POST …:test) copies the most recent conversation behind the card into a new sandbox session, up to the customer message the finding points at. It then opens the debug console with Trying: title. To start from a different conversation, use Try the change here on its evidence row. Each turn runs the live revision with the change made in memory. Nothing is saved: no revision, no draft, and no change to the card. Confirm writes is on, so any write that would reach a real system waits for you. If no conversation behind the card can be copied, you get a fresh sandbox session and write the message yourself.
Test on past conversations
Test the change (POST …:compare) replays each conversation twice, once on the deployed revision (Now (vN)) and once with the change (With the change). Each replay has the same history the agent had. A judge compares the two replies without being told which is which. The two replies are labelled A and B in an order drawn for each case. The judge runs on the agent's Evaluation model, which you set in Studio under Evaluation metrics. If none is set, the judge uses the agent's own model. A test starts two runs:
- Its own conversations: the conversations behind the card (up to 10 by default), from the turn the finding points at. The judge asks whether each reply avoids the problem.
- Other recent conversations: the tab replays 10 (the API default is 20, and the maximum for either run is 50), from where each conversation ended. The judge asks which reply is better, to catch harm elsewhere.
Each case comes out as Better with the change, Worse with the change, or No real difference. A case that couldn't be replayed or judged is counted as not compared. The card leads with a one-line verdict, such as Fixes the problem in 6 of 8 conversations or Fixes the problem, but 1 other conversation got worse. Below it are the tallies and the replies side by side. In the API, the runs are evaluation runs at …/evaluationRuns/{id}. summary.pairwise counts the wins, and score is the change's share of the decided cases, from 0 to 100.
Replays are isolated and recorded as test runs. Any tool that could change something is simulated. A connection with no saved sandbox mode is read through, so the agent's reads reach live data as it is now while everything else is simulated and rolled back. With the API, you can route a connection yourself with sandboxRouting (SIMULATED or READ_THROUGH). The card keeps only its latest test. That test is marked Out of date once another revision goes live or the change is revised. Saved replies expire when their conversation does.
Check the worse ones again
Replies vary from run to run, and the judge has to pick a side, so a single "worse" may be noise. Check the worse ones again (POST …:recheck, with the card's comparison.id as comparisonId) replays each worse case twice more on both sides. A case counts as worse only if the change lost at least two of its three verdicts. Each test can be checked once, or again if the check ended without a verdict.
Revise a change
POST …:revise is API only. It sends the cases a recheck confirmed back to the reviewer. Pass recheckId (the card's comparison.recheck.id), plus excludedCases and a note if you want. The reviewer answers with one of three outcomes:
REVISED: a rewritten change. It's tested straight away on the card's conversations, on every case shown to the reviewer, and on recent conversations it hasn't seen. If that test can't start,retestNotStartedsays why.KEPT: the change didn't cause the worse replies.WITHDRAW_SUGGESTED: the harm comes from the very thing the card asks for.
Whatever the answer, the card stays open and records it in revisions. A change can be rewritten at most twice. Revising runs on the reviewer model and the review budget, even when reviews are off, and often takes about half a minute.
Approve
Approve opens Make this change as vN. The dialog repeats what the last test found. Approve and deploy then does the following:
- Applies the change to the live revision. The live revision must still be the one you approved against (
expectedBaseRevision). - Validates the result the way a publish would.
- Saves it as a new revision with the note "Improvement: title".
- Deploys that revision to all of the agent's traffic.
If another revision goes live in the meantime, the new revision is saved but not deployed, the card stays open, and the response has deployed: false. Undo redeploys the revision the change replaced (POST …/revisions/{previousRevision}:deploy) and opens the card again. Teammates with unsaved drafts on the old revision need to update them before they publish, as after any deploy. See Versioning.
Approve is refused when the change can't be made safely. That happens if a tool still needs a connection, there's more than one account the tool could run on (attach it in the editor), the agent has more than one agent step, or the card or live revision changed since you opened it (409). If the live agent already has the change, the card closes.
Dismiss
Dismiss asks for a reason. Every later review of this agent sees the reason, so the reviewer learns what your admins don't want.
| Reason | What happens next |
|---|---|
| Not relevant | Raised again only if it keeps happening (five more conversations) |
| Wrong | Never suggested again |
| Already handled | Raised again only if it keeps happening. Mark done on an advice card uses this reason. |
| Other | The reviewer is told why it was dismissed. Add the explanation in note (API only, up to 1,000 characters). |
Weekly digest
Each week, from 06:00 UTC on Monday, Application admins who receive the Agent Improvements alert get a summary by email. The alert is on by default. The email covers each agent with reviews on and at least one waiting card. For each agent, it lists up to five waiting changes, what went live or was verified that week, how many conversations were reviewed or skipped because the budget ran out, how many questions went to the knowledge base, and what reviewing cost. See Observability → Alerts.
API
All endpoints live under /v1/tenants/{tenant}/agents/{agent}, and all of them require the caller to be an Application Admin. :approve and :dismiss require the card's version as you read it. If the card changed since then, you get a 409. Any change to a managed agent also gets a 409.
| Endpoint | Scope | Does |
|---|---|---|
GET · PUT …/reviewSettings | agents:read · agents:write | Read or replace the review settings |
POST …:review · GET …/conversationReview | agents:write · agents:read | Review a conversation now · read how it was reviewed |
GET …/recommendations · GET …/recommendations/{recommendation} | agents:read | List cards (filter with status; with no filter, everything except candidates) · one card and its evidence |
POST …/recommendations/{recommendation}:test | agents:test | Copy a conversation into a sandbox session to try the change |
POST …:compare · POST …:recheck | agents:test | Test on past conversations · replay the worse cases again |
POST …:revise | agents:write | Send confirmed worse cases back to the reviewer |
POST …:approve | revisions:write + revisions:deploy | Create and deploy the new revision |
POST …:dismiss | agents:write | Dismiss with a reason |
GET …/reviewFindings | agents:read | List findings, optionally by disposition |
Next steps
- Inline Judge: score live replies, and provide the quality-check signals that send conversations to review first
- Evaluations: test a revision against a fixed dataset
- Knowledge Bases: where knowledge-gap drafts are reviewed and published
- Versioning: revisions, deploys, and rollback
- Sandbox Mode: how replays and tries keep away from production