The difficult part may be verification itself.
Lyell and Coiera reviewed automation-bias studies and found that problems were not confined to multitasking: some single-task settings involved substantial verification complexity. Their analysis points to cognitive load as a research direction, with important limitations in the underlying evidence. Read the review.
Our application to patient access is a testable question: can a staff member establish whether a claimed action happened without reconstructing the workflow across inaccessible systems? The review did not test that particular question.
Inspect what the reviewer can actually see.
A reassuring summary and a completed-status label may come from the same unverified assertion. Giving a human both does not necessarily add independent evidence. Ask which authoritative record, event or receiving-team acknowledgment is available to that person.
A practical oversight walkthrough
In an authorized test environment, present a request whose summary and authoritative result disagree. Observe whether the responsible reviewer notices, locates the relevant evidence, understands the discrepancy and can take the agreed corrective action.
Record the permissions, systems, steps, time and assistance needed. A test with a specialist engineer beside the reviewer may not represent the ordinary operating arrangement.
Questions for operations, quality and security.
- What triggers a review, including requests that never produce a success message?
- Can the reviewer access the authoritative evidence during the relevant shift?
- What happens when the evidence is incomplete or conflicting?
- Can the reviewer recover the request or obtain acceptance from the correct team?
- Does the correction preserve a usable route for the legitimate patient or proxy?
- Have these steps been observed with the current configuration and staffing assumptions?
Interpret automation research carefully.
In Gaube and colleagues’ experimental study, inaccurate advice reduced diagnostic accuracy regardless of whether it was labeled as AI or human advice. The advice itself was generated by human experts. That makes it useful evidence about advice and reliance, not a measured error rate for a deployed AI product. Read the study.
The assessment opportunity is to inspect the actual oversight arrangement and its evidence. It is not to presume staff will fail, or to promise that additional training alone resolves the problem.
Sources & evidence limits.
Published research informs the questions. The application to your configured workflow requires its own evidence.
- Automation bias and verification complexity
Systematic review of 40 studies, six in healthcare. Heterogeneous tasks and limited controlled comparisons constrain generalization.
- Automation bias: frequency, mediators and mitigators
Systematic review across domains. Supports investigating reliance and workload; does not estimate current agent failure rates.
- Do as AI say
Experimental chest-X-ray task with expert-generated advice labeled AI or human. Not a patient-access or live-agent study.