Uncertain authority
Pause a new consequential write. Retain an owned assisted request and tell the user what remains pending.
THE METHOD
An answer, an attempted tool call and a committed record change are different observations. Our assessment keeps those distinctions visible.
Correlate the proposed action, permission decision, connector invocation, authoritative state before and after, and outbound communication. If the scheduling record is unavailable, report the committed outcome as unverified.
Pause a new consequential write. Retain an owned assisted request and tell the user what remains pending.
Reconcile authoritative state before retrying, declaring failure or attempting compensation. Exercise lost responses, duplicate requests and concurrent staff changes.
Apply the customer’s queue and escalation rules with a named owner and truthful status. Do not call an unowned callback request a completed handoff.
For rescheduling, test the customer’s rule for preserving the original arrangement while finding a replacement. Keep cancellation-only as a distinct, legitimate patient choice. Recovery must reconcile records, messages and staff work—not just restore a row.
Test legitimate patients and proxies, correction of ambiguous requests, supported languages and relevant accessibility paths. Operational owners supply the service limits and care context; the assessor does not invent clinical deadlines.
When automation is restricted, inspect the assisted route, pending queue and ownership. Measure offered requests and unresolved outcomes as well as successful completions. A system that refuses everyone has not demonstrated good patient access.
Separate speech recognition, speaker verification and permission to act. Specify the real call path, supported languages, transport, codecs, devices and noise conditions. A text prompt or laboratory audio result does not establish performance on the deployed phone path.
If speaker biometrics are present, agree thresholds and false-accept/false-reject measurement with appropriate samples. Do not infer demographic fairness or real-world attack feasibility from a small test set.
Trace untrusted messages, referral text and tool results into proposed capabilities. Test enforcement even when the model proposes the wrong action: approval binding, control-service timeouts, retries, background jobs and changed tool permissions. Assess MCP or other protocols only where they are deployed.
Check object-level authorization, account recovery and patient/proxy boundaries across the in-scope paths. These failures can exist without AI. OWASP’s object authorization guidance describes the underlying class.
Record the approved rule, environment, configuration versions, test method, repeated-trial count, evidence references, observed effects and limits. Report depth separately from breadth. More variations through one path do not automatically mean wider workflow coverage.
Use restricted evidence references and synthetic identifiers. Do not require private model reasoning or full clinical transcripts. Agree evidence handling before collection and preserve the customer’s responsibility for risk acceptance.
Retest the failure and legitimate counterpart after correction. Changes to the model, policy, connector, identity integration or workflow can invalidate prior conclusions and trigger a new scope decision.
Primary links checked 6 September 2026. These sources inform assessment design; they are not certifications, government endorsements or evidence that most hospitals are exploitable.
BEFORE THE NEXT ROLLOUT