THE METHOD

Follow the request.
Prove the outcome.

An answer, an attempted tool call and a committed record change are different observations. Our assessment keeps those distinctions visible.

Who can do what, for whom?

Begin with the customer’s rule: requester, represented patient, current proxy relationship, selected record, permitted operation and required confirmation. Map the system that enforces each part at execution.

Test stale sessions, revoked delegates, incorrect record references and changes after approval. Pair each prohibited action with an authorized counterpart. A recognized voice or valid login does not establish permission to change every appointment.

HHS documented social engineering of healthcare help desks in April 2024, including impersonation used to manipulate access recovery. That alert supports examining identity and recovery boundaries; it is not a measured patient-access failure rate. HHS HC3 alert.

What actually changed?

Correlate the proposed action, permission decision, connector invocation, authoritative state before and after, and outbound communication. If the scheduling record is unavailable, report the committed outcome as unverified.

01

Uncertain authority

Pause a new consequential write. Retain an owned assisted request and tell the user what remains pending.

02

Uncertain commit

Reconcile authoritative state before retrying, declaring failure or attempting compensation. Exercise lost responses, duplicate requests and concurrent staff changes.

03

Unavailable capacity

Apply the customer’s queue and escalation rules with a named owner and truthful status. Do not call an unowned callback request a completed handoff.

For rescheduling, test the customer’s rule for preserving the original arrangement while finding a replacement. Keep cancellation-only as a distinct, legitimate patient choice. Recovery must reconcile records, messages and staff work—not just restore a row.

A safeguard must leave a usable route.

Test legitimate patients and proxies, correction of ambiguous requests, supported languages and relevant accessibility paths. Operational owners supply the service limits and care context; the assessor does not invent clinical deadlines.

When automation is restricted, inspect the assisted route, pending queue and ownership. Measure offered requests and unresolved outcomes as well as successful completions. A system that refuses everyone has not demonstrated good patient access.

Test the architecture you actually use.

Voice agents and IVAs

Separate speech recognition, speaker verification and permission to act. Specify the real call path, supported languages, transport, codecs, devices and noise conditions. A text prompt or laboratory audio result does not establish performance on the deployed phone path.

If speaker biometrics are present, agree thresholds and false-accept/false-reject measurement with appropriate samples. Do not infer demographic fairness or real-world attack feasibility from a small test set.

Chat, agents and connected tools

Trace untrusted messages, referral text and tool results into proposed capabilities. Test enforcement even when the model proposes the wrong action: approval binding, control-service timeouts, retries, background jobs and changed tool permissions. Assess MCP or other protocols only where they are deployed.

Websites, portals and APIs

Check object-level authorization, account recovery and patient/proxy boundaries across the in-scope paths. These failures can exist without AI. OWASP’s object authorization guidance describes the underlying class.

Make the conclusion reproducible—and bounded.

Record the approved rule, environment, configuration versions, test method, repeated-trial count, evidence references, observed effects and limits. Report depth separately from breadth. More variations through one path do not automatically mean wider workflow coverage.

Use restricted evidence references and synthetic identifiers. Do not require private model reasoning or full clinical transcripts. Agree evidence handling before collection and preserve the customer’s responsibility for risk acceptance.

Retest the failure and legitimate counterpart after correction. Changes to the model, policy, connector, identity integration or workflow can invalidate prior conclusions and trigger a new scope decision.

Primary foundations

Primary links checked 6 September 2026. These sources inform assessment design; they are not certifications, government endorsements or evidence that most hospitals are exploitable.

BEFORE THE NEXT ROLLOUT

Start with the action
you need to trust.

Discuss your workflow