The evidence behind AI-assisted charting.
A growing body of independent, peer-reviewed research supports AI-assisted charting — and points directly to the design choices that make Scribe Mutual different.
The problem is evident.
In a time-and-motion study across four specialties, physicians spent nearly two hours on the EHR and desk work for every hour of direct patient care, plus another one to two hours of after-hours “pajama time” on the computer each night.1 EHR logs from primary care tell the same story: clinicians now split their day roughly evenly between patients and “desktop medicine.”2
That burden is one of the most consistently identified drivers of clinician burnout, with documentation and clerical load, cognitive overload, and time demands named again and again in the literature.3 4
And the notes this produces aren't necessarily better. Copy-paste and copy-forward are used by 66–90% of clinicians, fueling “note bloat,” internal inconsistency, and error propagation — one analysis linked copied documentation to 2.6% of diagnostic errors serious enough to send patients back for unplanned care.5
The goal isn't just faster notes. It's notes that are accurate, current, and drafted fresh from what actually happened in the room.
AI-assisted documentation already shows results.
In one of the largest real-world deployments to date, 7,260 physicians at The Permanente Medical Group used ambient AI documentation across roughly 2.5 million patient visits, and the health system estimated about 15,800 hours of documentation time saved — while reporting better patient–physician interaction and higher clinician satisfaction.6
~15,800 hrs
Documentation time saved across ~2.5M visits at The Permanente Medical Group6
51.9% → 38.8%
Drop in clinicians reporting burnout after 30 days, across six health systems7
6.2 → 5.3 min
Time in notes per appointment, with sharply lower mental demand8
Every serious review adds the same caveat: AI drafts can contain errors, omissions, and hallucinations, so diligent clinician oversight is essential.9 That caveat isn't a footnote for us — it's the design principle behind everything below.
Evidence for what makes Scribe Mutual different.
Ambient transcription is becoming a commodity. The hard, safety-critical work happens after the words are captured — and that's where our design choices are grounded in the research.
A problem list you can actually trust
Problem lists are frequently incomplete and out of date. A ten-organization study found completeness averaged 78% and ranged from 60% to 99% — a “global data integrity problem that could compromise quality of care and put patients at risk.”10 NLP that pulls problems straight from the encounter has long been proposed to keep them current.11 Scribe Mutual builds every problem from what was said in the visit, with provenance back to the transcript.
Negation- and family-history-aware
Naive extraction quietly drags ruled-out findings and family history onto the active list. Purpose-built negation and family-history detection reaches F-scores around 0.91, materially improving accuracy.11 12 13 Scribe Mutual keeps “no chest pain” and “mother had diabetes” off the active problem list.
Structure first, prose second
Language models can fabricate or drop clinical facts; under adversarial conditions, hallucination rates have been measured at 50–82%.14 15 But a disciplined pipeline can push hallucination and omission below human note-taking rates (1.47% and 3.45% in one framework).16 Scribe Mutual extracts structured facts in a separate pass from narrative writing, so data stays traceable — not invented by a paragraph generator.
The clinician signs — always
Every credible review of AI documentation stresses that notes need clinician oversight before they enter the record.9 Scribe Mutual notes stay editable drafts until you sign them. Nothing is auto-signed, billed, or shared from a draft.
Coding help that assists, not decides
AI-assisted ICD-10 coding has been shown to raise coder accuracy (F1 from 0.83 to 0.92) and support faster review in real hospital workflows.17 18 Scribe Mutual suggests ICD-10 and E/M codes tied to the visit's content and flags specificity gaps — you decide.
FHIR-native, open to your patients
Patients overwhelmingly value reading their notes: 98% call portal access a good idea, most say it helps them understand and follow their care, and traditionally underserved patients report the greatest benefit.19 20 21 Scribe Mutual is built on FHIR R4 with US Core profiles and releases signed notes to the patient portal immediately.
References
Peer-reviewed sources cited above. Links resolve to each paper's abstract and source; deployment figures reflect the health systems' own published reports.
- Sinsky C, et al. Allocation of Physician Time in Ambulatory Practice: A Time and Motion Study in 4 Specialties. Annals of Internal Medicine. 2016.
- Tai-Seale M, et al. Electronic Health Record Logs Indicate That Physicians Split Time Evenly Between Seeing Patients and Desktop Medicine. Health Affairs. 2017.
- Budd J. Burnout Related to Electronic Health Record Use in Primary Care. Journal of Primary Care & Community Health. 2023.
- Moy AJ, et al. Measurement of clinical documentation burden among physicians and nurses using electronic health records: a scoping review. JAMIA. 2021.
- Tsou AY, et al. Safe Practices for Copy and Paste in the EHR. Applied Clinical Informatics. 2017.
- Tierney AA, et al. Ambient Artificial Intelligence Scribes to Alleviate the Burden of Clinical Documentation (and Learnings after 1 Year and over 2.5 Million Uses). NEJM Catalyst. 2024–2025. Kaiser Permanente / The Permanente Medical Group.
- Olson KD, et al. Use of Ambient AI Scribes to Reduce Administrative Burden and Professional Burnout. JAMA Network Open. 2025.
- Stults CD, et al. Evaluation of an Ambient Artificial Intelligence Documentation Platform for Clinicians. JAMA Network Open. 2025.
- Leung TI, et al. AI Scribes in Health Care: Balancing Transformative Potential With Responsible Integration. JMIR Medical Informatics. 2025.
- Wright A, et al. Problem list completeness in electronic health records: a multi-site study and assessment of success factors. International Journal of Medical Informatics. 2015.
- Meystre S, et al. Natural language processing to extract medical problems from electronic clinical documents: Performance evaluation. Journal of Biomedical Informatics. 2006.
- Garcelon N, et al. Improving a full-text search engine: the importance of negation detection and family history context to identify cases in a biomedical data warehouse. JAMIA. 2017.
- Bill R, et al. Automated Extraction of Family History Information from Clinical Notes. AMIA Annual Symposium Proceedings. 2014.
- Omar M, et al. Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support. Communications Medicine. 2025.
- Shah SV. Accuracy, Consistency, and Hallucination of Large Language Models When Analyzing Unstructured Clinical Notes in Electronic Medical Records. JAMA Network Open. 2024.
- Asgari E, et al. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. npj Digital Medicine. 2025.
- Chen PF, et al. Automatic ICD-10 Coding and Training System: Deep Neural Network Based on Supervised Learning. JMIR Medical Informatics. 2021.
- Dai HJ, et al. Evaluating a Natural Language Processing–Driven, AI-Assisted ICD-10-CM Coding System for Diagnosis Related Groups in a Real Hospital Environment. JMIR. 2024.
- Walker J, et al. OpenNotes After 7 Years: Patient Experiences With Ongoing Access to Their Clinicians' Outpatient Visit Notes. Journal of Medical Internet Research. 2019.
- DesRoches CM, et al. Patients Managing Medications and Reading Their Visit Notes: A Survey of OpenNotes Participants. Annals of Internal Medicine. 2019.
- Bell SK, et al. Tackling Ambulatory Safety Risks Through Patient Engagement: What 10,000 Patients and Families Say After Reading Visit Notes. Journal of Patient Safety. 2018.
Want to see how this works in your workflow?
We'd love to talk through your documentation before handing over credentials. Beta approvals usually take a few business days, and you'll hear back from a real person — most likely the founder.