The application code is unchanged. The model behind an API alias has moved to a new version. Responses still look fluent, but the agent now chooses a different tool more often. The release process needs to recognize that the system has changed even though the repository has not.
Review the behaviour a change can affect
An approval belongs to a particular use, configuration, and operating boundary. Changing a model, prompt, retrieval source, tool schema, or credential can disturb a different part of that boundary. Requiring the same evidence for every change wastes effort while making consequential differences harder to see.
| Change | Question to resolve | Evidence to collect |
|---|---|---|
| Model or system prompt | Does task behaviour still meet expectations? | Comparison on representative tasks and difficult cases |
| Retrieval corpus | Has the information boundary changed? | Source ownership, access and relevance checks |
| Tool contract | Can calls now have different effects? | Integration tests and dependency review |
| Permission expansion | Who authorized the added reach? | Specific approval and denial tests |
| Output destination | Who can now receive the result? | Recipient, exposure and retention review |
These categories can overlap. A tool change may also alter data access, and a prompt revision may expose a previously unused capability. The release owner should describe the expected impact before choosing tests. That gives reviewers something concrete to challenge.
Use a baseline you can reproduce
Keep the current configuration and a small set of representative cases. Compare the candidate with the version actually running, using the same inputs where feasible. Examine incorrect actions, unnecessary escalation, missed refusals, and changed tool use alongside answer quality, cost, and latency.
Some services do not let the customer pin or restore every underlying component. Record that limitation in the operating plan and decide how to detect changes. A rollback promise is weak if the supplier can no longer provide the previous model or if the changed workflow has already written to external systems.
A rollback has two parts
Restoring software returns the system to a previous configuration. Repairing its effects addresses records, messages, permissions, or transactions already changed. Plan both where the agent can act. A canary release limits exposure, but the team still needs to identify which users and operations encountered the candidate version.
For federal systems covered by Canada’s Directive on Automated Decision-Making, functionality or scope changes can also trigger an update to the published AIA. That obligation has a defined scope. Other organizations should establish their own reapproval triggers according to the consequences of their use.
Watch the behaviour that justified release
Choose monitoring that reflects the change being approved. If a release changes tool selection, inspect tool-choice errors. If it expands a data source, watch access failures and inappropriate retrieval. A general uptime graph cannot establish that the revised behaviour remains acceptable.
The release record should make the next comparison easier: identifiable versions, observed results, known limitations, and a decision owner. Over time, this creates a history of what the organization actually approved and why, including the cases where it decided to wait.
Selected primary sources
Open the primary-source pages used to verify the claims summarized here.