Compares a recommendation card's change with the live agent across past conversations.
Where `:test` forks one conversation for you to read, this replays many and has a judge compare the replies. Each conversation is replayed from the customer message at a turn, with the history the agent had, twice: on the deployed release, and on the release with the change made in memory. A judge that is not told which is which says which reply is better, or that neither is. Two runs are started: <list type="bullet"><item>`cardRun` — the conversations behind the card, at the turn their finding points at, judged on whether the reply avoids what the review found;</item><item>`otherRun` — the agent's other recent conversations, at the turn where they ended, judged on which reply is better: what else the change touches.</item></list> Both are evaluation runs: read them at `evaluationRuns/{id}` until `COMPLETED`. The run's `summary.pairwise` counts the cases each side won, and its `score` is the changed agent's share of the decided ones, 0–100; `samples` holds each conversation's two replies and the judge's reason, and expires when the conversation does. The card keeps the latest comparison as `comparison`, so it can show the result to whoever opens it next; nothing about the agent changes, and nothing leaves the replays: tools that could change something are simulated. Agent admins only. No connection needs a sandbox set up first. One without a saved sandbox mode is read through for the replays: the reads the agent is known to make reach the live service, so the replies rest on real data, and everything else is simulated and rolled back when the replays end. Choose per connection with `sandboxRouting`; nothing is saved on the connection.
Authorization
Bearer JWT Authorization header using the Bearer scheme. Enter 'Bearer' [space] and then your token.
In: header
Path Parameters
The agent ID.
The recommendation ID.
uuidThe tenant identifier
How many conversations of each kind; omit either for its default.
How many of the conversations behind the card to replay, newest first, each at the turn its finding points at. Default Fruxon.Model.AgentReviews.CompareRecommendationRequest.DefaultCardConversations; 0 skips them.
int32How many of the agent's other recent conversations to replay, each at the turn where it ended — the ones the change was not aimed at. Default Fruxon.Model.AgentReviews.CompareRecommendationRequest.DefaultOtherConversations; 0 skips them.
int32How to route a connection the agent uses during the replays, by connection id: SIMULATED or
READ_THROUGH. Nothing is saved on the connection. A connection left out keeps its saved sandbox mode
when it has one, and is otherwise read through: reads the agent is known to make reach the live service,
everything else is simulated and rolled back when the replays end.
Response Body
curl -X POST "https://api.fruxon.com/v1/tenants/string/agents/string/recommendations/497f6eca-6276-4993-bfeb-53cbbbba6f08:compare" \ -H "Content-Type: application/json" \ -d '{}'{
"id": "00000000-0000-0000-0000-000000000000",
"baseRevision": 0,
"cardRun": {
"id": "00000000-0000-0000-0000-000000000000",
"agentId": "string",
"kind": "DATASET",
"datasetId": "00000000-0000-0000-0000-000000000000",
"candidateRevision": 0,
"candidateBaseRevision": 0,
"candidateProposalId": "00000000-0000-0000-0000-000000000000",
"candidateSnapshotSeqNo": 0,
"deployedRevision": 0,
"status": "PENDING",
"score": 0,
"deploymentRecommendation": "string",
"summary": {
"totalSamples": 0,
"successfulEvaluations": 0,
"failedEvaluations": 0,
"averageLlmScore": 0,
"scoreDistribution": {},
"assessmentDistribution": {},
"topConcerns": [
"string"
],
"keyImprovements": [
"string"
],
"pairwise": {
"cases": 0,
"candidateWins": 0,
"baselineWins": 0,
"ties": 0,
"failed": 0,
"decided": 0,
"minDecidedCases": null,
"score": null
},
"runtime": {
"totalCandidateMs": 0,
"totalBaseMs": 0,
"averageCandidateMs": 0,
"averageBaseMs": 0,
"stdevCandidateMs": 0
},
"cost": {
"totalCandidate": 0,
"totalBase": 0,
"averageCandidate": 0,
"averageBase": 0
}
},
"errorMessage": "string",
"createdAt": 0,
"modifiedAt": 0,
"startedAt": 0,
"completedAt": 0,
"createdBy": "string"
},
"otherRun": {
"id": "00000000-0000-0000-0000-000000000000",
"agentId": "string",
"kind": "DATASET",
"datasetId": "00000000-0000-0000-0000-000000000000",
"candidateRevision": 0,
"candidateBaseRevision": 0,
"candidateProposalId": "00000000-0000-0000-0000-000000000000",
"candidateSnapshotSeqNo": 0,
"deployedRevision": 0,
"status": "PENDING",
"score": 0,
"deploymentRecommendation": "string",
"summary": {
"totalSamples": 0,
"successfulEvaluations": 0,
"failedEvaluations": 0,
"averageLlmScore": 0,
"scoreDistribution": {},
"assessmentDistribution": {},
"topConcerns": [
"string"
],
"keyImprovements": [
"string"
],
"pairwise": {
"cases": 0,
"candidateWins": 0,
"baselineWins": 0,
"ties": 0,
"failed": 0,
"decided": 0,
"minDecidedCases": null,
"score": null
},
"runtime": {
"totalCandidateMs": 0,
"totalBaseMs": 0,
"averageCandidateMs": 0,
"averageBaseMs": 0,
"stdevCandidateMs": 0
},
"cost": {
"totalCandidate": 0,
"totalBase": 0,
"averageCandidate": 0,
"averageBase": 0
}
},
"errorMessage": "string",
"createdAt": 0,
"modifiedAt": 0,
"startedAt": 0,
"completedAt": 0,
"createdBy": "string"
},
"regressionRun": {
"id": "00000000-0000-0000-0000-000000000000",
"agentId": "string",
"kind": "DATASET",
"datasetId": "00000000-0000-0000-0000-000000000000",
"candidateRevision": 0,
"candidateBaseRevision": 0,
"candidateProposalId": "00000000-0000-0000-0000-000000000000",
"candidateSnapshotSeqNo": 0,
"deployedRevision": 0,
"status": "PENDING",
"score": 0,
"deploymentRecommendation": "string",
"summary": {
"totalSamples": 0,
"successfulEvaluations": 0,
"failedEvaluations": 0,
"averageLlmScore": 0,
"scoreDistribution": {},
"assessmentDistribution": {},
"topConcerns": [
"string"
],
"keyImprovements": [
"string"
],
"pairwise": {
"cases": 0,
"candidateWins": 0,
"baselineWins": 0,
"ties": 0,
"failed": 0,
"decided": 0,
"minDecidedCases": null,
"score": null
},
"runtime": {
"totalCandidateMs": 0,
"totalBaseMs": 0,
"averageCandidateMs": 0,
"averageBaseMs": 0,
"stdevCandidateMs": 0
},
"cost": {
"totalCandidate": 0,
"totalBase": 0,
"averageCandidate": 0,
"averageBase": 0
}
},
"errorMessage": "string",
"createdAt": 0,
"modifiedAt": 0,
"startedAt": 0,
"completedAt": 0,
"createdBy": "string"
}
}{
"type": "string",
"title": "string",
"status": 0,
"detail": "string",
"instance": "string",
"property1": null,
"property2": null
}{
"type": "string",
"title": "string",
"status": 0,
"detail": "string",
"instance": "string",
"property1": null,
"property2": null
}{
"type": "string",
"title": "string",
"status": 0,
"detail": "string",
"instance": "string",
"property1": null,
"property2": null
}{
"type": "string",
"title": "string",
"status": 0,
"detail": "string",
"instance": "string",
"property1": null,
"property2": null
}