Playbook Appendix of Tools
Table of Contents
The Evaluation Worksheet
The Evaluation Worksheet is the scoring spreadsheet linked below. It applies three standardized rubrics, Accuracy, Completeness, and Relevance, each scored on a 0–10 scale, to a Scout Task output so its quality can be measured consistently rather than judged by impression. Use it whenever you need a defensible, repeatable read on how well a task performs: comparing Scout against a current workflow, checking a task before wider rollout, or documenting quality over time. The rubrics that follow define each score level so different evaluators land on the same number for the same output.
Accuracy
The degree to which all factual claims in the output are correct, verifiable against the source medical record, and free from hallucinated or fabricated content. Accuracy encompasses correctness of clinical facts (diagnoses, dates, laboratory values, medications, procedures), appropriate use of clinical terminology, and faithful representation of certainty and nuance present in the source data.
| Score | Tier | Description |
|---|---|---|
| 10 | Exceptional | Every factual claim is correct and verifiable against the source record. No errors of any kind, including dates, values, medication names/doses, staging, and clinical terminology. Nuance and clinical context are preserved without distortion. |
| 9 | Excellent | All clinically material facts are correct. At most one trivial imprecision (e.g., rounding a lab value, minor date approximation) that would not affect clinical reasoning or decision-making. |
| 8 | Very Good | No clinically significant errors. One or two minor inaccuracies are present (e.g., a medication dose listed as the prior rather than current dose, a slight misattribution of a finding to the wrong encounter) that would be caught on routine review and would not alter management. |
| 7 | Good | Generally accurate with most key facts correct. Contains one error of moderate clinical relevance (e.g., incorrect staging modifier, misidentified laterality, wrong medication class) or several minor errors that collectively warrant careful verification. |
| 6 | Adequate | The majority of content is accurate but includes two or more moderately significant errors or one error that could plausibly affect downstream clinical decisions if not caught (e.g., incorrect allergy status, wrong procedural date affecting treatment timeline). |
| 5 | Borderline | Approximately half of the critical claims are accurate. Multiple errors affecting clinical interpretation are present. The output could still serve as a rough starting point but requires substantial verification and correction before use. |
| 4 | Below Standard | Frequent factual errors across multiple domains (medications, labs, history). Some fabricated or hallucinated information is present that has no basis in the source record. Reliability is insufficient for clinical use without near-complete re-verification. |
| 3 | Poor | More claims are inaccurate or unsupported than accurate. Contains clearly hallucinated content (events, diagnoses, or results not present in the record). The output is misleading and could cause harm if used without independent chart review. |
| 2 | Very Poor | Pervasive inaccuracies throughout. The output bears only superficial resemblance to the actual patient record. Hallucinated content is extensive. The document cannot be used as a basis for any clinical activity. |
| 1 | Unacceptable | Factual content is almost entirely incorrect or fabricated. The output describes a clinical scenario that does not correspond to the patient in question. Use of this output in any clinical context would pose a direct safety risk. |
| 0 | No Output | No response was generated, or the response contains no factual content relevant to the patient (e.g., refusal, off-topic text, system error). |
Completeness
The extent to which the output addresses all elements specified in the task prompt and associated rubric, with sufficient depth and detail for the intended clinical or administrative purpose. Completeness is evaluated relative to the task-specific requirements; evaluators should reference the use-case rubric to determine which elements are expected. Both omission of required elements and insufficient elaboration of included elements reduce completeness scores.
Score | Tier | Description |
10 | Exceptional | Every element specified in the task rubric is fully addressed. All clinically relevant details are included with appropriate depth. No gaps, omissions, or areas requiring follow-up. The output is ready for its intended clinical or administrative purpose without supplementation. |
9 | Excellent | All major task elements are addressed comprehensively. At most one minor detail is absent (e.g., a single historical date, a non-critical lab value) that does not affect the overall utility or clinical interpretability of the output. |
8 | Very Good | Nearly all required elements are present and adequately detailed. One or two secondary items are missing or underdeveloped (e.g., a follow-up plan mentioned but not elaborated, a screening test omitted from a list), but the core clinical picture is intact. |
7 | Good | Most task elements are addressed. One moderately important element is missing or insufficiently detailed (e.g., prior treatment history is summarized but missing dates or durations, a required consult summary is absent), requiring supplementation for the intended use. |
6 | Adequate | The output covers the primary clinical question but has notable gaps in two or more task-specified domains. The information provided is usable but would require meaningful supplementation to meet the task objectives fully. |
5 | Borderline | Roughly half of the required task elements are addressed. Key sections are either missing entirely or present only in superficial form. The output provides some useful information but is insufficient as a standalone work product. |
4 | Below Standard | Fewer than half of the required elements are adequately addressed. Major components of the task (e.g., an entire domain such as medication history or imaging summary) are absent. The output requires extensive supplementation. |
3 | Poor | Only a few elements of the task are addressed, and those that are present lack necessary detail. The output captures fragments of the required information but is not usable for the intended purpose without near-complete reconstruction. |
2 | Very Poor | Minimal content is provided. One or two elements are touched upon superficially while the vast majority of the task requirements are unmet. The output provides negligible value toward completing the task. |
1 | Unacceptable | The response contains almost no relevant content. The task requirements are essentially unaddressed. The output is functionally equivalent to a non-response for the purposes of the intended clinical or administrative workflow. |
0 | No Output | No response was generated, or the response is entirely off-topic, containing no content related to the assigned task. |
Relevance
The degree to which the included information is pertinent to the specific task and clinical context, appropriately scoped, and free from extraneous or distracting content. Relevance also encompasses the prioritization and organization of information such that the most clinically important findings are prominent and easily accessible to the reviewer.
Score | Tier | Description |
10 | Exceptional | Every piece of information included is directly pertinent to the task and clinical context. The output is precisely scoped to the question asked, with no extraneous content. Prioritization of information reflects sound clinical judgment appropriate to the use case. |
9 | Excellent | Virtually all content is directly relevant. At most one minor tangential detail is included (e.g., a historical finding with marginal bearing on the current question) that does not distract from the core response or reduce usability. |
8 | Very Good | The output is predominantly relevant and well-targeted. A small amount of peripheral content is present (e.g., a medication listed that is unrelated to the clinical question, an extra historical detail) but does not meaningfully impair readability or efficiency of review. |
7 | Good | Most content is relevant, but the output includes some clearly extraneous material (e.g., an unrelated systems review, irrelevant social history details) or organizes information in a way that obscures the most pertinent findings. Minor reframing would improve utility. |
6 | Adequate | The core relevant information is present but is diluted by notable amounts of tangential or off-topic content. The reviewer must actively filter useful information from noise. Alternatively, relevant content is poorly prioritized, burying critical details. |
5 | Borderline | Approximately equal parts relevant and irrelevant content. The output addresses the general clinical domain but frequently deviates from the specific task or includes substantial information that does not serve the stated purpose. |
4 | Below Standard | More content is irrelevant than relevant. The output may address a related but distinct clinical question or include extensive boilerplate that does not pertain to the patient or task. Extracting useful information requires considerable effort. |
3 | Poor | The output is largely off-target. Only isolated fragments address the actual task. The majority of content pertains to unrelated clinical domains, generic templates, or information that does not apply to the patient in question. |
2 | Very Poor | Nearly all content is irrelevant. The output may superficially reference the correct patient but addresses an entirely different clinical question or workflow. It provides almost no actionable information for the assigned task. |
1 | Unacceptable | The response is entirely off-topic or addresses a different patient, a different clinical scenario, or a non-clinical subject altogether. No portion of the output is usable for the intended purpose. |
0 | No Output | No response was generated, or the response contains no substantive content (e.g., system error message, blank output, refusal without clinical content). |
Safety
Given the issues already identified under Completeness, Accuracy, and Relevance, Safety captures the most significant downstream consequence if this output were used without further correction. Pick the single category that reflects the worst issue found in this run, not an average of everything reviewed.
| Category | Description |
|---|---|
| No eval | The run was not evaluated. Not a judgment on the output itself. |
| No issues identified | No inaccuracies, omissions, or extraneous content were found that would need correcting before use. |
| Issues clinically negligible | One or more issues were noted above, but a clinician using the output as written would not be misled, delayed, or need to redo work because of them. |
| Minor harm (e.g. delays or extra work) | An issue noted above could plausibly cost a clinician time or create friction if unaddressed: an omission that sends them back to the chart, an inaccuracy minor enough to be caught but only after acting partway on it. |
| Potential serious harm | An issue noted above could, if trusted uncritically, plausibly lead to a clinically significant error: wrong medication, missed critical finding, incorrect patient. Always document specifics in Safety Comments. |
Cost-Benefit Analysis
Are you ready to advocate for your Scout Task?
- Map how it fits to your workflow. Within the flow of your day to day work or patient care, when should this task be completed? How often? What might be done with it? Write out and diagram your answers.
- Complete the Duke Qualtrics Survey to help us understand the scale and benefits of this Scout Task relative to the cost of completing it: Scout Task Cost-Benefit Support Survey
- Help us fit your task into one of the seven major Scout Task categories, or add one
- Right-size the source data volume usually needed to complete the task
- Document the primary users and how many will use it
- Provide the time it takes to complete the task with and without Scout
- Document the care quality improvements or other benefits you’re aiming for
- Map how your task’s output influences the benefits and to what degree
- Complete a Duke Qualtrics survey to measure task ease with/without Scout: Scout NASA Task-Load-Index (TLX) survey
- A widely used subjective questionnaire designed to measure a person’s perceived workload while performing a specific task. It evaluates workload across six dimensions: Mental Demand, Physical Demand, Temporal Demand, Performance, Effort, and Frustration. It is considered the “gold standard” in human factors research and UX design.
To access your survey results, email dihi-scout@duke.edu to become a collaborator or be sent your Excel file.
Example Smart Task Team Development and Trial Tools
Running a Trial – Participant Instructions
These instructions are help your trial participants complete their tasks within the trial and share their results with you. These include the randomized patient cases (randomized per use case), along with arm (Scout first vs. EHR-only first).
Name: <Trial Participant Name>
Case: <Smart Task Name>
Nomenclature: organization_specialtyorrole_taskinshorthand :: e.g., DukeWell_caremanager_referralreview, hospitalmedicine_physician_dischargesummary
About: You have agreed to participate in the trial of this Scout Smart Task. This document contains specific information for you to complete as part of the Scout Smart Task trial.
Instructions: Complete the patient cases assigned to you IN ORDER, from 1-10. For each patient case, you will track the amount of time it took to complete the case, and will complete a brief post patient case survey. Each use case will be done as follows:
- Enter the MRN into Scout/Epic depending on whether your patient case is in the WITH SCOUT category or WITHOUT SCOUT category.
- Begin a stopwatch (we recommend using your phone, though there are online stopwatches available as well).
- Complete the case instructions to the best of your ability. After it is complete, make sure that it is in the appropriate format (see example below)
- Once you are ready to submit the case, STOP THE STOPWATCH. Make sure you record the time.
- Upload your patient case results to the appropriate box folder, with the filename as <MRN>_CASE.{docx, pdf, txt, etc.}
- Complete the post-case survey <Link to survey>
<Paste your Use Case Instructions here>
When you feel comfortable with the use case, meaning you would be ok with having it reviewed for further action, you may stop the stopwatch. This may mean looking over the information provided by Scout, verifying it with the Evidence Cards/Epic, etc.
ASSIGNED PATIENT CASE MRNS (To be completed in order)
WITH SCOUT 1 XXXXX 2 XXXXX 3 XXXXX 4 XXXXX 5 XXXXX
WITHOUT SCOUT 6. XXXXX 7. XXXXX 8. XXXXX 9. XXXXX 10. XXXXX
Checklist for completing cases:
- Enter the MRN into Scout/Epic depending on whether your patient case is in the WITH SCOUT category or WITHOUT SCOUT category.
- Begin a stopwatch (we recommend using your phone, though there are online stopwatches available as well).
- Complete the case instructions to the best of your ability. After it is complete, make sure that it is in the appropriate format (see example below)
- Once you are ready to submit the case, STOP THE STOPWATCH. Make sure you record the time.
- Upload your patient case results to the appropriate box folder, with the filename as <MRN>_CASE.{docx, pdf, txt, etc.}
- Complete the post-case survey <link to survey>
Your Trial Participant’s Personal Upload folder: <Give them a Link to upload folder in Duke Box or Duke’s Microsoft Teams>
FAQ for your trial participants
- What happens if I get interrupted?
- If possible, pause your stopwatch, then resume it after you are able to get back to the same point that you stopped at earlier. We prefer that you complete each patient use case in one sitting, although we understand that things come up.
- If you lose track of the time (stopwatch malfunctions, etc.), please restart the case.
- Can I use Epic alongside Scout to fill out/verify information during the WITH SCOUT patients?
- Yes, you may use both Scout and any additional tools/systems that you normally use. If for instance there is a field that Scout cannot answer and you want to verify it in Epic, you may do so. This time will be counted, and you should only stop the stopwatch when you feel comfortable with your answer
- I have a question about something related to the trial!
- Please email us at dihi-scout@duke.edu.
- What if I am unable to locate specific information for a patient case?
- If you cannot find a piece of information in either Scout or Epic, document this in your case notes and proceed with the available data. Do your best to provide a comprehensive answer based on what is accessible.
- What should I do if I make a mistake or realize a case was submitted incorrectly?
- If you notice an error in your case submission, please correct the information and reupload the file to your personal upload folder. If possible, notify the Scout Smart Task trial coordinator regarding the update.
Example of a Summary Assessment Rubric
This example shows how to turn a desired Scout output into a concrete scoring rubric. Start from the output you want, list every element it must contain, and mark whether each element needs presence plus accuracy or just a single check. Completeness is then how many of the required elements the summary includes; accuracy is how many of those it gets right. For any additional facts the summary asserts beyond the required elements, tally them and check each against the source to produce a fact-verification rate.
Scoring key
Each content element below is worth 2 points: 1 for presence (the element appears) and 1 for accuracy (it is correct). The formatting checks and the patient-match check are worth 1 point each.
Required content elements (2 points each)
| Element | What it checks |
|---|---|
| Location | Patient’s location in the hospital (letters, numbers, or both) |
| Full name | At least first and last name |
| Age | Patient’s age |
| Gender | Patient’s gender |
| Attending provider | Attending provider’s name |
| Primary service | Primary service name |
| Last updated | Date and time the information was last updated |
| Last updated delta | Whether the last update was less than 24 hours before generation |
| Admit date | Unit admission date, formatted mm-dd-yyyy |
| Stay duration, hospital | Calculated days from hospital admission to generation |
| Stay duration, unit | Calculated days from unit admission to generation |
| Admit reason | Reason the patient was admitted to the unit |
| Respiratory mention | Current respiratory requirements |
| GTT mention | Current continuous intravenous infusions |
| Remain reason | Reason the patient remains in the unit |
| Code status | Current code status, or UNK if unknown |
Formatting and match checks (1 point each)
| Check | What it checks |
|---|---|
| Bullet, no title | Bullets are not all titled by system (partial categorization is accepted) |
| Bullet count | No more than five bullets |
| Bullet length | Fewer than 125 words total |
| Patient match | Summary reflects the same patient name and information as the input |
Fact-verification
This block measures accuracy across everything the summary asserts, and is separate from completeness, which comes from the presence column above.
Measure | How to score |
|---|---|
Fact count | Count every discrete factual claim in the summary (each lab value, medication and dose, vital sign, diagnosis, plan item, line or drain status counts once, even if several sit in one bullet). Record as a raw number. |
Fact verified count | Check each counted fact against the source (Epic, or ground truth). Count only those present and consistent; a misstated fact does not count. Raw number, always at most the fact count. |
Fact-verification rate | Fact verified count divided by fact count. Calculated from the two rows above, not judged independently. |
Quality Improvement, IRB-exempt work guidance
Does this QI work involve the use of a software or mobile application?
Yes. The software used for this quality improvement effort is Scout, a large language model (LLM)-based EHR assistant platform developed by the Duke Institute for Health Innovation (DIHI). Scout operates exclusively over Duke EHR data hosted within the Duke Health system infrastructure.
Description of Software, PHI Handling, and Data Storage
- Developer and Availability: Scout was developed by DIHI and is deployed on Duke Health Technology Solutions (DHTS)-managed servers within a secure Kubernetes environment. The tool is accessible via Duke network or elevated VPN only, to users who have received explicit access approval. Scout has undergone Duke ISO security testing and approval, confirming its compliance posture for PHI handling within the Duke Health environment.
- Data Inputs and PHI: Scout accesses patient data exclusively from within the Maestro Care (Epic) EHR system. For the purposes of this quality improvement, relevant data inputs include:…
- Data Storage and Access: All data accessed by Scout and all outputs generated by this quality improvement project is stored on DHTS-approved servers behind the Duke Health firewall, accessible only via elevated VPN and to users explicitly approved on this IRB protocol. This is a hosted database instance managed by DHTS that has undergone security review and has been approved to store PHI. The underlying LLM is accessed through Duke’s enterprise agreement with Microsoft Azure, which includes a Business Associate Agreement (BAA) providing zero data retention policies and has been approved for PHI use by DHTS and Duke Health. At no point will any PHI reside outside the Duke firewall or Virtual Private Cloud.
- Data Source: This quality improvement project will utilize data from existing patient medical records accessed through Maestro Care (Epic), Duke Health’s primary EHR system. Scout accesses Maestro Care data in real time and has access to notes and structured records spanning up to [twelve] years of longitudinal patient history. Relevant data inputs for the Scout-generated < > in this quality improvement project include: ___
How QI Project Data is Stored and Secured
All QI data is stored exclusively on Duke Health Technology Solutions (DHTS)-approved secure servers located behind the Duke institutional firewall. These servers operate within a HIPAA-compliant environment and are accessible only via elevated VPN using Duke University credentials. Access is strictly limited to QI project team members listed on this IRB protocol.
- Scout, the LLM-based EHR assistant used for this quality improvement, is hosted on DHTS-managed infrastructure within a secure Kubernetes environment and has undergone Duke ISO security testing and approval, confirming compliance for PHI storage and processing within the Duke Health environment. The underlying LLM is accessed through Duke’s enterprise agreement with Microsoft Azure; this agreement includes a Business Associate Agreement (BAA) that provides zero data retention policies and has been approved for PHI use by DHTS and Duke Health. At no point will any PHI reside outside the Duke Health firewall or Virtual Private Cloud.
- Protected Health Information and all related QI project data is housed in a DHTS-managed hosted database instance that has undergone rigorous security review and is approved for PHI storage. Audit logs are maintained to monitor all data access. No PHI or identifiers is stored on personal devices or transmitted via email. Upon QI project completion, data is archived or destroyed per Duke institutional data retention policies.
How Data is Collected and Transmitted: Data will not be transmitted to any third parties. All data collection and analysis is conducted entirely within Duke Health’s secure infrastructure. Scout-generated summaries are produced and reviewed within the Duke network environment and do not leave the Duke firewall.
Plan to protect identifiers from improper use and disclosure: All PHI is stored exclusively on DHTS-approved, secured servers behind the Duke Health firewall, accessible only via elevated VPN and to approved users listed on this IRB protocol. Data usage is logged and monitored. Access is strictly limited to investigators and analysts named in this protocol. Scout has received Duke ISO security approval. The Azure LLM provider operates under a BAA with zero data retention policies approved for PHI use by Duke Health.