Advancing Italian biomedical information extraction with transformers-based models: Methodological insights and multicenter practical application

[Display omitted] The introduction of computerized medical records in hospitals has reduced burdensome activities like manual writing and information fetching. However, the data contained in medical records are still far underutilized, primarily because extracting data from unstructured textual medi...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	Journal of biomedical informatics 2023-12, Vol.148, p.104557-104557, Article 104557
Hauptverfasser:	Crema, Claudio, Buonocore, Tommaso Mario, Fostinelli, Silvia, Parimbelli, Enea, Verde, Federico, Fundarò, Cira, Manera, Marina, Ramusino, Matteo Cotta, Capelli, Marco, Costa, Alfredo, Binetti, Giuliano, Bellazzi, Riccardo, Redolfi, Alberto
Format:	Artikel
Sprache:	eng
Schlagworte:	Biomedical text mining Data Mining - methods Deep learning Electronic Health Records Humans Italy Language Language model Natural Language Processing Transformer
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	[Display omitted] The introduction of computerized medical records in hospitals has reduced burdensome activities like manual writing and information fetching. However, the data contained in medical records are still far underutilized, primarily because extracting data from unstructured textual medical records takes time and effort. Information Extraction, a subfield of Natural Language Processing, can help clinical practitioners overcome this limitation by using automated text-mining pipelines. In this work, we created the first Italian neuropsychiatric Named Entity Recognition dataset, PsyNIT, and used it to develop a Transformers-based model. Moreover, we collected and leveraged three external independent datasets to implement an effective multicenter model, with overall F1-score 84.77 %, Precision 83.16 %, Recall 86.44 %. The lessons learned are: (i) the crucial role of a consistent annotation process and (ii) a fine-tuning strategy that combines classical methods with a “low-resource” approach. This allowed us to establish methodological guidelines that pave the way for Natural Language Processing studies in less-resourced languages.
ISSN:	1532-0464 1532-0480
DOI:	10.1016/j.jbi.2023.104557