Responsible AI resourceMenuHomeModel policyDecisionsResearch & approachGet support
Leadership guide/Operations & assurance
What evidence justifies a live pilot or wider use?
Download the guide↓Editable working records→
Define success and unacceptable failure before testing. Evaluate the complete service in the conditions where people will use it, including human review, connected actions and recovery.
Responsible roles: Service owner and designated approver with independent challenge proportionate to consequence.
Leave with: A release decision supported by results and known limitations.
Choose realistic tasks, the existing service baseline and measurable outcomes. Include accuracy, unsupported answers, accessibility, harmful disparities, time to human help, privacy, cost and workload where relevant. Set thresholds locally with the people responsible for the consequences.
Cover ordinary use, ambiguous requests, missing or outdated information, relevant languages and accessibility needs. Include mistakes people are likely to make. Keep test data authorized, and document limitations in population or workflow coverage.
Test misleading external instructions, unauthorized data requests and attempts to exceed connected permissions. For agents, include wrong recipients, duplicate actions, partial completion, delegated activity, tool failure, spending limits and the ability to stop and reconcile effects.
Check that reviewers have enough information, time and authority to disagree. Walk through an appeal, service outage and correction with actual role holders. Confirm an accessible alternative and fallback for essential services.
Record results, failures, corrections, remaining uncertainty and the configuration tested. Apply the deployment-blocking conditions in this guide. A pilot decision names users, duration, limits, monitoring and stop conditions. Expansion requires pilot evidence and an explicit decision; material changes trigger retesting.
Apply these conditions before a live pilot, wider deployment or material expansion. A favorable average test result does not offset a failed mandatory safeguard. Keep any experiment isolated from affected people, protected records and production actions until the conditions are met.
Do not release when testing shows access to information or actions outside approved authority. Restrict the scope, fix the boundary and retest.
Do not release while required data-handling protections, processing authority or contractual commitments remain unresolved. Obtain evidence and the necessary approval for the actual use.
Do not release without a named owner, an effective way to stop or disconnect the service, and a person authorized and prepared to use it. Demonstrate the stop before live use.
Do not release an essential service without a tested fallback and a responsible operator. Include the effects of partial completion and recovery of affected records.
Resolve other critical security findings before release, or document a permissible, authorized and time-limited exception with demonstrated compensating safeguards. Exceptions cannot waive applicable requirements or the blocking conditions above.
For each applicable condition, record the evidence, result, reviewer and date. Use the release decision field in the existing working record. The designated authority records the approved scope; missing evidence leaves the affected condition open.
Illustrative campus example
An advising assistant answers common questions well but confidently gives outdated financial-aid deadlines. The team restricts those answers until authoritative sources, escalation and test results meet its release conditions. An average success rate cannot cancel an unacceptable high-consequence failure.
Use the editable companion to capture the decision in your institution’s own systems. These prompts are also available in the printable guide.
State the version, environment, intended users, baseline, thresholds and unacceptable failures.
Record representative and adversarial cases, expected outcomes, actual results and evidence locations.
Document human review, accessibility, fallback, corrections and remaining uncertainty.
Record each applicable deployment-blocking condition and evidence of closure, any permissible exception, the approver, scope, monitoring and retest triggers.
Model policy: Section 4 · Section 6 · Section 9 · Section 12.
HumanSkills implementation recommendations informed by the sources below. These voluntary resources are not legal requirements or a certification. Apply current institutional requirements and specialist review to the actual use.