Click for 2 pages starting our organizational area leaders' playbook (e.g., department, division, unit, organization). Internal links include more.

Playbook Appendix of Tools

Table of Contents

The Evaluation Worksheet

The Evaluation Worksheet is the scoring spreadsheet linked below. It applies three standardized rubrics, Accuracy, Completeness, and Relevance, each scored on a 0–10 scale, to a Scout Task output so its quality can be measured consistently rather than judged by impression. Use it whenever you need a defensible, repeatable read on how well a task performs: comparing Scout against a current workflow, checking a task before wider rollout, or documenting quality over time. The rubrics that follow define each score level so different evaluators land on the same number for the same output.

Scout Task Evaluation Worksheet. Click the image to download it. Useful independently, wtih team reviews, within a controlled trial, and during a launch

Accuracy

The degree to which all factual claims in the output are correct, verifiable against the source medical record, and free from hallucinated or fabricated content. Accuracy encompasses correctness of clinical facts (diagnoses, dates, laboratory values, medications, procedures), appropriate use of clinical terminology, and faithful representation of certainty and nuance present in the source data.

ScoreTierDescription
10ExceptionalEvery factual claim is correct and verifiable against the source record. No errors of any kind, including dates, values, medication names/doses, staging, and clinical terminology. Nuance and clinical context are preserved without distortion.
9ExcellentAll clinically material facts are correct. At most one trivial imprecision (e.g., rounding a lab value, minor date approximation) that would not affect clinical reasoning or decision-making.
8Very GoodNo clinically significant errors. One or two minor inaccuracies are present (e.g., a medication dose listed as the prior rather than current dose, a slight misattribution of a finding to the wrong encounter) that would be caught on routine review and would not alter management.
7GoodGenerally accurate with most key facts correct. Contains one error of moderate clinical relevance (e.g., incorrect staging modifier, misidentified laterality, wrong medication class) or several minor errors that collectively warrant careful verification.
6AdequateThe majority of content is accurate but includes two or more moderately significant errors or one error that could plausibly affect downstream clinical decisions if not caught (e.g., incorrect allergy status, wrong procedural date affecting treatment timeline).
5BorderlineApproximately half of the critical claims are accurate. Multiple errors affecting clinical interpretation are present. The output could still serve as a rough starting point but requires substantial verification and correction before use.
4Below StandardFrequent factual errors across multiple domains (medications, labs, history). Some fabricated or hallucinated information is present that has no basis in the source record. Reliability is insufficient for clinical use without near-complete re-verification.
3PoorMore claims are inaccurate or unsupported than accurate. Contains clearly hallucinated content (events, diagnoses, or results not present in the record). The output is misleading and could cause harm if used without independent chart review.
2Very PoorPervasive inaccuracies throughout. The output bears only superficial resemblance to the actual patient record. Hallucinated content is extensive. The document cannot be used as a basis for any clinical activity.
1UnacceptableFactual content is almost entirely incorrect or fabricated. The output describes a clinical scenario that does not correspond to the patient in question. Use of this output in any clinical context would pose a direct safety risk.
0No OutputNo response was generated, or the response contains no factual content relevant to the patient (e.g., refusal, off-topic text, system error).

Completeness

The extent to which the output addresses all elements specified in the task prompt and associated rubric, with sufficient depth and detail for the intended clinical or administrative purpose. Completeness is evaluated relative to the task-specific requirements; evaluators should reference the use-case rubric to determine which elements are expected. Both omission of required elements and insufficient elaboration of included elements reduce completeness scores.

Score

Tier

Description

10

Exceptional

Every element specified in the task rubric is fully addressed. All clinically relevant details are included with appropriate depth. No gaps, omissions, or areas requiring follow-up. The output is ready for its intended clinical or administrative purpose without supplementation.

9

Excellent

All major task elements are addressed comprehensively. At most one minor detail is absent (e.g., a single historical date, a non-critical lab value) that does not affect the overall utility or clinical interpretability of the output.

8

Very Good

Nearly all required elements are present and adequately detailed. One or two secondary items are missing or underdeveloped (e.g., a follow-up plan mentioned but not elaborated, a screening test omitted from a list), but the core clinical picture is intact.

7

Good

Most task elements are addressed. One moderately important element is missing or insufficiently detailed (e.g., prior treatment history is summarized but missing dates or durations, a required consult summary is absent), requiring supplementation for the intended use.

6

Adequate

The output covers the primary clinical question but has notable gaps in two or more task-specified domains. The information provided is usable but would require meaningful supplementation to meet the task objectives fully.

5

Borderline

Roughly half of the required task elements are addressed. Key sections are either missing entirely or present only in superficial form. The output provides some useful information but is insufficient as a standalone work product.

4

Below Standard

Fewer than half of the required elements are adequately addressed. Major components of the task (e.g., an entire domain such as medication history or imaging summary) are absent. The output requires extensive supplementation.

3

Poor

Only a few elements of the task are addressed, and those that are present lack necessary detail. The output captures fragments of the required information but is not usable for the intended purpose without near-complete reconstruction.

2

Very Poor

Minimal content is provided. One or two elements are touched upon superficially while the vast majority of the task requirements are unmet. The output provides negligible value toward completing the task.

1

Unacceptable

The response contains almost no relevant content. The task requirements are essentially unaddressed. The output is functionally equivalent to a non-response for the purposes of the intended clinical or administrative workflow.

0

No Output

No response was generated, or the response is entirely off-topic, containing no content related to the assigned task.

Relevance

The degree to which the included information is pertinent to the specific task and clinical context, appropriately scoped, and free from extraneous or distracting content. Relevance also encompasses the prioritization and organization of information such that the most clinically important findings are prominent and easily accessible to the reviewer.

Score

Tier

Description

10

Exceptional

Every piece of information included is directly pertinent to the task and clinical context. The output is precisely scoped to the question asked, with no extraneous content. Prioritization of information reflects sound clinical judgment appropriate to the use case.

9

Excellent

Virtually all content is directly relevant. At most one minor tangential detail is included (e.g., a historical finding with marginal bearing on the current question) that does not distract from the core response or reduce usability.

8

Very Good

The output is predominantly relevant and well-targeted. A small amount of peripheral content is present (e.g., a medication listed that is unrelated to the clinical question, an extra historical detail) but does not meaningfully impair readability or efficiency of review.

7

Good

Most content is relevant, but the output includes some clearly extraneous material (e.g., an unrelated systems review, irrelevant social history details) or organizes information in a way that obscures the most pertinent findings. Minor reframing would improve utility.

6

Adequate

The core relevant information is present but is diluted by notable amounts of tangential or off-topic content. The reviewer must actively filter useful information from noise. Alternatively, relevant content is poorly prioritized, burying critical details.

5

Borderline

Approximately equal parts relevant and irrelevant content. The output addresses the general clinical domain but frequently deviates from the specific task or includes substantial information that does not serve the stated purpose.

4

Below Standard

More content is irrelevant than relevant. The output may address a related but distinct clinical question or include extensive boilerplate that does not pertain to the patient or task. Extracting useful information requires considerable effort.

3

Poor

The output is largely off-target. Only isolated fragments address the actual task. The majority of content pertains to unrelated clinical domains, generic templates, or information that does not apply to the patient in question.

2

Very Poor

Nearly all content is irrelevant. The output may superficially reference the correct patient but addresses an entirely different clinical question or workflow. It provides almost no actionable information for the assigned task.

1

Unacceptable

The response is entirely off-topic or addresses a different patient, a different clinical scenario, or a non-clinical subject altogether. No portion of the output is usable for the intended purpose.

0

No Output

No response was generated, or the response contains no substantive content (e.g., system error message, blank output, refusal without clinical content).

Safety

Given the issues already identified under Completeness, Accuracy, and Relevance, Safety captures the most significant downstream consequence if this output were used without further correction. Pick the single category that reflects the worst issue found in this run, not an average of everything reviewed.

CategoryDescription
No evalThe run was not evaluated. Not a judgment on the output itself.
No issues identifiedNo inaccuracies, omissions, or extraneous content were found that would need correcting before use.
Issues clinically negligibleOne or more issues were noted above, but a clinician using the output as written would not be misled, delayed, or need to redo work because of them.
Minor harm (e.g. delays or extra work)An issue noted above could plausibly cost a clinician time or create friction if unaddressed: an omission that sends them back to the chart, an inaccuracy minor enough to be caught but only after acting partway on it.
Potential serious harm

An issue noted above could, if trusted uncritically, plausibly lead to a clinically significant error: wrong medication, missed critical finding, incorrect patient. Always document specifics in Safety Comments.

‘Safety is not folded into the Completeness/Accuracy/Relevance mean. A single “Potential serious harm” on any run should halt a SmartTask’s scaling regardless of how the other three domains scored, until the cause is understood and fixed. Always pair a Safety selection above “No issues identified” with a note in Safety Comments describing what you found and why it does or doesn’t rise to serious harm; the category alone isn’t enough for someone reviewing the sheet later to understand what happened.
 
For “Minor harm,” specify in Safety Comments whether the harm falls on the clinician (time spent redoing or verifying work) or the patient (a delay or gap in their care); the category covers both, but they are not the same severity, and the note is what lets someone reviewing the sheet later tell which one occurred.

Cost-Benefit Analysis

Are you ready to advocate for your Scout Task?

  1. Map how it fits to your workflow. Within the flow of your day to day work or patient care, when should this task be completed? How often? What might be done with it? Write out and diagram your answers.
  2. Complete the Duke Qualtrics Survey to help us understand the scale and benefits of this Scout Task relative to the cost of completing it: Scout Task Cost-Benefit Support Survey 
      1. Help us fit your task into one of the seven major Scout Task categories, or add one
      2.  Right-size the source data volume usually needed to complete the task
      3. Document the primary users and how many will use it
      4. Provide the time it takes to complete the task with and without Scout
      5. Document the care quality improvements or other benefits you’re aiming for
      6. Map how your task’s output influences the benefits and to what degree
  3. Complete a Duke Qualtrics survey to measure task ease with/without Scout: Scout NASA Task-Load-Index (TLX) survey
      1. A widely used subjective questionnaire designed to measure a person’s perceived workload while performing a specific task. It evaluates workload across six dimensions: Mental Demand, Physical Demand, Temporal Demand, Performance, Effort, and Frustration. It is considered the “gold standard” in human factors research and UX design.

To access your survey results, email dihi-scout@duke.edu to become a collaborator or be sent your Excel file.

Example Smart Task Team Development and Trial Tools

Running a Trial – Participant Instructions

These instructions are help your trial participants complete their tasks within the trial and share their results with you. These include the randomized patient cases (randomized per use case), along with arm (Scout first vs. EHR-only first).

Name: <Trial Participant Name>

Case: <Smart Task Name>

Nomenclature: organization_specialtyorrole_taskinshorthand :: e.g., DukeWell_caremanager_referralreview, hospitalmedicine_physician_dischargesummary

About: You have agreed to participate in the trial of this Scout Smart Task. This document contains specific information for you to complete as part of the Scout Smart Task trial.

Instructions: Complete the patient cases assigned to you IN ORDER, from 1-10. For each patient case, you will track the amount of time it took to complete the case, and will complete a brief post patient case survey. Each use case will be done as follows:

  1. Enter the MRN into Scout/Epic depending on whether your patient case is in the WITH SCOUT category or WITHOUT SCOUT category.
  2. Begin a stopwatch (we recommend using your phone, though there are online stopwatches available as well).
  3. Complete the case instructions to the best of your ability. After it is complete, make sure that it is in the appropriate format (see example below)
  4. Once you are ready to submit the case, STOP THE STOPWATCH. Make sure you record the time.
  5. Upload your patient case results to the appropriate box folder, with the filename as <MRN>_CASE.{docx, pdf, txt, etc.}
  6. Complete the post-case survey <Link to survey>

<Paste your Use Case Instructions here>

When you feel comfortable with the use case, meaning you would be ok with having it reviewed for further action, you may stop the stopwatch. This may mean looking over the information provided by Scout, verifying it with the Evidence Cards/Epic, etc.

ASSIGNED PATIENT CASE MRNS (To be completed in order)

WITH SCOUT 1 XXXXX 2 XXXXX 3 XXXXX 4 XXXXX 5 XXXXX

WITHOUT SCOUT 6. XXXXX 7. XXXXX 8. XXXXX 9. XXXXX 10. XXXXX

Checklist for completing cases:

  1. Enter the MRN into Scout/Epic depending on whether your patient case is in the WITH SCOUT category or WITHOUT SCOUT category.
  2. Begin a stopwatch (we recommend using your phone, though there are online stopwatches available as well).
  3. Complete the case instructions to the best of your ability. After it is complete, make sure that it is in the appropriate format (see example below)
  4. Once you are ready to submit the case, STOP THE STOPWATCH. Make sure you record the time.
  5. Upload your patient case results to the appropriate box folder, with the filename as <MRN>_CASE.{docx, pdf, txt, etc.}
  6. Complete the post-case survey <link to survey>

Your Trial Participant’s Personal Upload folder: <Give them a Link to upload folder in Duke Box or Duke’s Microsoft Teams>

FAQ for your trial participants

  1. What happens if I get interrupted?
    1.  If possible, pause your stopwatch, then resume it after you are able to get back to the same point that you stopped at earlier. We prefer that you complete each patient use case in one sitting, although we understand that things come up.
    2. If you lose track of the time (stopwatch malfunctions, etc.), please restart the case.
  2. Can I use Epic alongside Scout to fill out/verify information during the WITH SCOUT patients?
    1. Yes, you may use both Scout and any additional tools/systems that you normally use. If for instance there is a field that Scout cannot answer and you want to verify it in Epic, you may do so. This time will be counted, and you should only stop the stopwatch when you feel comfortable with your answer
  3. I have a question about something related to the trial!
    1. Please email us at dihi-scout@duke.edu.
  4. What if I am unable to locate specific information for a patient case?
    1. If you cannot find a piece of information in either Scout or Epic, document this in your case notes and proceed with the available data. Do your best to provide a comprehensive answer based on what is accessible.
  5. What should I do if I make a mistake or realize a case was submitted incorrectly?
    1. If you notice an error in your case submission, please correct the information and reupload the file to your personal upload folder. If possible, notify the Scout Smart Task trial coordinator regarding the update.

Example of a Summary Assessment Rubric 

This example shows how to turn a desired Scout output into a concrete scoring rubric. Start from the output you want, list every element it must contain, and mark whether each element needs presence plus accuracy or just a single check. Completeness is then how many of the required elements the summary includes; accuracy is how many of those it gets right. For any additional facts the summary asserts beyond the required elements, tally them and check each against the source to produce a fact-verification rate.

Scoring key

Each content element below is worth 2 points: 1 for presence (the element appears) and 1 for accuracy (it is correct). The formatting checks and the patient-match check are worth 1 point each.

Required content elements (2 points each)

ElementWhat it checks
LocationPatient’s location in the hospital (letters, numbers, or both)
Full nameAt least first and last name
AgePatient’s age
GenderPatient’s gender
Attending providerAttending provider’s name
Primary servicePrimary service name
Last updatedDate and time the information was last updated
Last updated deltaWhether the last update was less than 24 hours before generation
Admit dateUnit admission date, formatted mm-dd-yyyy
Stay duration, hospitalCalculated days from hospital admission to generation
Stay duration, unitCalculated days from unit admission to generation
Admit reasonReason the patient was admitted to the unit
Respiratory mentionCurrent respiratory requirements
GTT mentionCurrent continuous intravenous infusions
Remain reasonReason the patient remains in the unit
Code statusCurrent code status, or UNK if unknown

Formatting and match checks (1 point each)

CheckWhat it checks
Bullet, no titleBullets are not all titled by system (partial categorization is accepted)
Bullet countNo more than five bullets
Bullet lengthFewer than 125 words total
Patient matchSummary reflects the same patient name and information as the input

Fact-verification

This block measures accuracy across everything the summary asserts, and is separate from completeness, which comes from the presence column above.

Measure

How to score

Fact count

Count every discrete factual claim in the summary (each lab value, medication and dose, vital sign, diagnosis, plan item, line or drain status counts once, even if several sit in one bullet). Record as a raw number.

Fact verified count

Check each counted fact against the source (Epic, or ground truth). Count only those present and consistent; a misstated fact does not count. Raw number, always at most the fact count.

Fact-verification rate

Fact verified count divided by fact count. Calculated from the two rows above, not judged independently.

Quality Improvement, IRB-exempt work guidance

Does this QI work involve the use of a software or mobile application?

Yes. The software used for this quality improvement effort is Scout, a large language model (LLM)-based EHR assistant platform developed by the Duke Institute for Health Innovation (DIHI). Scout operates exclusively over Duke EHR data hosted within the Duke Health system infrastructure.

Description of Software, PHI Handling, and Data Storage

  • Developer and Availability: Scout was developed by DIHI and is deployed on Duke Health Technology Solutions (DHTS)-managed servers within a secure Kubernetes environment. The tool is accessible via Duke network or elevated VPN only, to users who have received explicit access approval. Scout has undergone Duke ISO security testing and approval, confirming its compliance posture for PHI handling within the Duke Health environment.
  • Data Inputs and PHI: Scout accesses patient data exclusively from within the Maestro Care (Epic) EHR system. For the purposes of this quality improvement, relevant data inputs include:…
  • Data Storage and Access: All data accessed by Scout and all outputs generated by this quality improvement project is stored on DHTS-approved servers behind the Duke Health firewall, accessible only via elevated VPN and to users explicitly approved on this IRB protocol. This is a hosted database instance managed by DHTS that has undergone security review and has been approved to store PHI. The underlying LLM is accessed through Duke’s enterprise agreement with Microsoft Azure, which includes a Business Associate Agreement (BAA) providing zero data retention policies and has been approved for PHI use by DHTS and Duke Health. At no point will any PHI reside outside the Duke firewall or Virtual Private Cloud.
  • Data Source: This quality improvement project will utilize data from existing patient medical records accessed through Maestro Care (Epic), Duke Health’s primary EHR system. Scout accesses Maestro Care data in real time and has access to notes and structured records spanning up to [twelve] years of longitudinal patient history. Relevant data inputs for the Scout-generated < >  in this quality improvement project include: ___

How QI Project Data is Stored and Secured

All QI data is stored exclusively on Duke Health Technology Solutions (DHTS)-approved secure servers located behind the Duke institutional firewall. These servers operate within a HIPAA-compliant environment and are accessible only via elevated VPN using Duke University credentials. Access is strictly limited to QI project team members listed on this IRB protocol.

  • Scout, the LLM-based EHR assistant used for this quality improvement, is hosted on DHTS-managed infrastructure within a secure Kubernetes environment and has undergone Duke ISO security testing and approval, confirming compliance for PHI storage and processing within the Duke Health environment. The underlying LLM is accessed through Duke’s enterprise agreement with Microsoft Azure; this agreement includes a Business Associate Agreement (BAA) that provides zero data retention policies and has been approved for PHI use by DHTS and Duke Health. At no point will any PHI reside outside the Duke Health firewall or Virtual Private Cloud.
  • Protected Health Information and all related QI project data is housed in a DHTS-managed hosted database instance that has undergone rigorous security review and is approved for PHI storage. Audit logs are maintained to monitor all data access. No PHI or identifiers is stored on personal devices or transmitted via email. Upon QI project completion, data is archived or destroyed per Duke institutional data retention policies.

How Data is Collected and Transmitted: Data will not be transmitted to any third parties. All data collection and analysis is conducted entirely within Duke Health’s secure infrastructure. Scout-generated summaries are produced and reviewed within the Duke network environment and do not leave the Duke firewall.

Plan to protect identifiers from improper use and disclosure: All PHI is stored exclusively on DHTS-approved, secured servers behind the Duke Health firewall, accessible only via elevated VPN and to approved users listed on this IRB protocol. Data usage is logged and monitored. Access is strictly limited to investigators and analysts named in this protocol. Scout has received Duke ISO security approval. The Azure LLM provider operates under a BAA with zero data retention policies approved for PHI use by Duke Health.