AI implementation
AI pilot acceptance checklist: proceed, revise or stop
Accept an AI pilot only for a defined next stage, using evidence from agreed test cases rather than a polished demonstration. For a document-to-draft workflow, check the extracted work, the human effort and the controls separately. Require evidence that missing information stays unresolved, approval applies to the exact draft revision, prohibited actions are blocked and uncertain handoffs are reconciled before retrying. If a required test fails or remains unassessed, hold that next stage, revise the scope or stop.
This is Amulet's recommended acceptance method for Australian business owners and operating leaders. It is not a certification standard, legal advice, an accuracy target or a report of tests we have performed. Every proposed case below starts as Not assessed. The decision is whether to authorise a bounded next stage. It is not a declaration that the system is generally safe or ready for unrestricted use.
Set the boundary before testing
Amulet's public document-to-draft demonstration describes fixed-format synthetic requests, deterministic parsing and validation, an approval step and a simulated handoff. It is not a live AI model or an ERP connector. Its scope excludes PDF/OCR accuracy, authenticated approver identity and production security, monitoring and support.1
Use it to discuss where work should stop, not as evidence that your documents, users or destination system have passed these tests. The schedule below does not report results from running the public demonstration.
Before testing, have the workflow owner complete a short test brief:
- Job and next stage sought: name the incoming document, the required draft and the action being considered. Distinguish preparing drafts for review from sending them to a business system.
- Version under test: record the input pack, reference data, model or parser, prompt, rules, settings and destination environment. Keep the versions needed to reproduce a result.
- People: name the test operator, business approver, system administrator and destination owner. Say who can accept each result and who can stop the pilot. A supplier's demonstration is not the business owner's acceptance.
- Expected work: define required fields, permitted corrections, unresolved-item handling and which actions are forbidden until approval. Decide which errors block progression and what review workload the team can accept before seeing the results.
- Safe test conditions: agree permitted data and access, a test destination and how to contain deliberate failures. Use synthetic material until representative business records are authorised. Do not inject failures into a live process merely to complete this checklist.
Keep cases used to tune the workflow separate from the acceptance set. Include the document formats and exceptions relevant to the proposed next stage. Record exclusions: testing clean text is not a reason to mark scanned-document handling as assessed.
Keep the simpler option in the test
Give the same required output and approval boundary to the manual process, relevant existing-software configuration and any proposed custom workflow. Ask your administrator to verify the native option under your actual product, plan, permissions and settings. Do not assume a feature exists or is included.
If a manual draft template and named reviewer meet the need, retain them. If existing software passes the required cases with acceptable effort, prefer that over commissioning equivalent custom work. Where explicit rules can do the job, test those before adding AI. If none of the options can be assessed because source records or responsibilities are unclear, fix the process first.
The partner selection guide covers supplier questions, and the implementation cost worksheet covers the resulting cost and ownership commitments. This schedule is for deciding what the pilot has demonstrated.
Measure extraction and review work separately
For an AI-assisted option, ask the business approver to set the expected result for each acceptance document before the run. Record each required field as correct, incorrect or unresolved, including whether its source reference supports the proposed value. Separately record whether the workflow surfaced each exception it was expected to catch.
Do not combine cosmetic differences, consequential quantity errors and missed exceptions into a single pass score. Set the acceptable error types and limits for this workflow before testing; this guide supplies no universal accuracy percentage. Record the case mix and the number of documents assessed beside any result. Keep synthetic control checks separate from observations on authorised representative documents.
Have the intended operator record the whole task: preparation, checking, corrections, approval, exception handling and recovery. Include rejected and abandoned work rather than timing only successful drafts. Compare that effort with the manual or native option on the same job, and ask the manager responsible for the team to accept the workload.
A well-extracted draft does not earn permission to skip an approval test, and a passing control case is not an extraction benchmark.
Run a case-based control schedule
Copy each case into the record below and assign actual names to the roles. The suggested runners and accepting roles are recommendations, not people already appointed to your pilot.
For every case, save the input and draft revisions, action attempted, actor and permission context, timestamps, workflow status and destination evidence. Where an action must be blocked, check both the workflow record and the agreed destination for unintended changes. Do not accept a success banner or an error message alone as the result.
A case may be marked Pass only when the agreed behaviour is observed and its evidence is saved. Mark an observed breach Fail. Missing evidence, an unavailable test environment or an unresolved outcome stays Not assessed, with a reason. If a case is outside the proposed scope, record the owner's exclusion and limit the next stage accordingly; do not count it as a pass.
Missing or conflicting input
Run: The operator submits a request missing a required quantity, then a separate request with contradictory delivery information. The business approver accepts the result.
Expected evidence: The draft identifies the affected fields and source passages, leaves the uncertainty visible and routes it to the named exception owner. Check that neither case creates a released destination record. The reviewer records the correction or rejection and its reason.
Fail or retest: Fail if a required value is silently guessed, a conflict disappears or release proceeds without resolution. After a fix, rerun both cases and a complete request to check that ordinary work can still reach review.
Initial result: Not assessed.
Approval attached to the draft revision
Run: The operator prepares a complete draft and tries to release it before approval. The authorised approver then approves it; the operator changes a material field and tries to release the changed version. The business approver and destination owner accept the result.
Expected evidence: Neither the unapproved draft nor the changed draft is released. The approval record identifies the actor and exact revision or content fingerprint. A new approval is required for the changed draft, and the permitted destination record matches that newly approved version. Inspect the approval history and destination record together.
Fail or retest: Fail if a general "approved" flag carries across changed content or if an earlier approval permits a different payload. After a fix, rerun unapproved, unchanged-approved and changed-after-approval paths.
Initial result: Not assessed.
Rejection and denied permission
Run: The business approver rejects a draft. The operator attempts its release. Separately, the administrator uses an agreed test identity without the required permission to try approval and access to a restricted test record. The business approver and administrator accept the results.
Expected evidence: Rejection remains visible with a reason and a route back for correction. The rejected draft is not released. The unauthorised identity cannot approve or retrieve the restricted content. Save the actual identity and permission context, denial records and destination checks, not just the role shown on a screen.
Fail or retest: Treat unauthorised access or release as a blocker, even if other cases pass. Fix the boundary and rerun both denied and permitted paths. A simulated role selector does not assess authenticated identity.
Initial result: Not assessed.
Duplicate submission and genuine amendment
Run: The operator resubmits the same approved request, then submits changed content using the same request identifier. Include overlapping submissions if that is possible in the proposed workflow. The destination owner accepts the result.
Expected evidence: An identical repeat refers back to the existing outcome rather than creating another business record. Changed content is not silently discarded as a duplicate or released under the old approval. It follows the agreed amendment or new-request process, with a traceable relationship to the original and fresh approval where required. Inspect destination record identifiers and contents across the attempts.
Fail or retest: Fail on an extra business record, a lost amendment or reuse of approval for changed content. After a fix, rerun the repeat and amendment paths, including the overlapping case where applicable.
Initial result: Not assessed.
Known failure before the destination accepts work
Run: In the agreed test environment, the administrator causes a handoff to fail before a destination record is created. The operator follows the recovery instructions; the destination owner accepts the result.
Expected evidence: Save evidence of the pre-creation failure and a destination check confirming no record exists within the tested scope. The workflow must not report completion. It should retain the approved revision and route the work to a named recovery owner. After the fault is removed, the controlled retry should create only the intended record.
Fail or retest: Fail on false completion, lost work or a duplicate after recovery. If the team cannot establish whether a record exists, do not label the event a known failure; use the unknown-outcome case next. Rerun interruption and recovery after a fix.
Initial result: Not assessed.
Unknown handoff outcome
Run: The administrator arranges a test in which the destination accepts the record but the sender does not receive confirmation. The operator reconciles the outcome before taking further action. The destination owner accepts the result.
Expected evidence: The workflow exposes an unresolved outcome rather than reporting success or assuming nothing happened. The operator checks the destination using the request reference and approved content, records what was found and resolves the status without creating a second record. Save evidence from both sides of the handoff.
Fail or retest: Fail if a blind retry creates a duplicate or the workflow declares completion without evidence. If the destination cannot be checked, keep the outcome unresolved and escalate; do not mark the test as passed. Rerun the lost-confirmation case and the known-failure case after a fix.
Initial result: Not assessed.
Recovery by the intended operator
Run: Give the intended operator a held case and the written instructions, without the person or supplier who built the workflow directing each step. Include a manual fallback and return to normal processing. The workflow owner and destination owner accept the result.
Expected evidence: The operator identifies the outstanding work, knows who may resolve it and either completes the authorised recovery or escalates with processing still held. Record manual actions so resuming the workflow does not repeat them. Check the final destination state and record the staff effort, assistance and unresolved items.
Fail or retest: Fail if the instructions require an undocumented approval bypass, work disappears or resumption repeats a manual action. If the builder must intervene, record that dependency and assess it against the agreed operating model. Revise the instructions or scope and rerun with the intended operator.
Initial result: Not assessed.
Keep one evidence record per case and attempt
Copy this record for each option and each run. Preserve failed attempts when retesting; a later pass should not erase what changed.
| Record field | Entry to complete |
|---|---|
| Option, case and attempt reference | |
| Stage sought and test environment | |
| Input, reference-data and workflow versions | |
| Named runner and accepting person | |
| Expected result and agreed pass condition | |
| Input and draft revision; approval actor and revision | |
| Action, identity/permission context and timestamps | |
| Observed output and unresolved fields or exceptions | |
| Destination lookup, record identifier and content evidence, or evidence supporting no change | |
| Evidence references accessible to the accepting person | |
| Handling effort, corrections and assistance required | |
| Result | Not assessed |
| Reason, blocker and responsible owner | |
| Fix, affected cases and retest evidence | |
| Accepting person's decision and date |
Keep a test's status separate from the transaction's status. An observed duplicate is a failed test; a handoff whose destination state cannot be established remains unresolved. Neither should disappear into an overall success rate.
Decide the next stage
Have the workflow owner make a written decision using the case records and the separate extraction and workload assessment:
- Proceed within a boundary: required cases have passed, the business approver accepts the output quality, and the operating team accepts the review and recovery work. Name the permitted inputs, users, destination, actions and limits for that stage. Keep unassessed formats and integrations outside it.
- Revise and retest: a gap has a credible remedy. Name the change, owner and evidence needed. Rerun affected cases and related approval, duplicate and recovery paths rather than retesting only the screen that changed.
- Stop or keep the simpler option: the pilot requires prohibited actions, unacceptable errors, unavailable permissions or more review work than the team accepts, or the manual/native option meets the need without the proposed build.
Record the decision, tested version, evidence references, approved boundary, exclusions and unresolved items, accountable owner, stop conditions and next review date. Leave the decision as Not assessed until the owner completes it. Agree who can pause processing and what evidence is required to resume. A pass on this schedule is not a guarantee about untested work or later changes.
Next step
Amulet's published capability includes defining a workflow's trigger, approved inputs, expected output, exception path, human roles, integration assumptions and acceptance criteria.2 That describes the published scope, not delivery results or an included connector for your systems.
Discuss your first AI workflow if you want to check fit. Bring a non-confidential description of the document, the draft it should produce, who may approve it and the handoff you have not yet resolved. Keep business records and confidential attachments out of the initial enquiry.
Sources
Sources
- Document-to-draft controls demonstration | Amulet AI
Retrieved
- Capabilities | Amulet AI
Retrieved
A practical next step
Put AI to work with the operating boundary visible.
Approvals, evidence and the rollout path should be mapped to the real workflow.
Discuss your first AI workflow