ConTEXTual Net: A Multimodal Vision-Language Model for Segmentation of Pneumothorax

Radiology narrative reports often describe characteristics of a patient's disease, including its location, size, and shape. Motivated by the recent success of multimodal learning, we hypothesized that this descriptive text could guide medical image analysis algorithms. We proposed a novel visio...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	Journal of digital imaging 2024-08, Vol.37 (4), p.1652-1663
Hauptverfasser:	Huemann, Zachary, Tie, Xin, Hu, Junjie, Bradshaw, Tyler J
Format:	Artikel
Sprache:	eng
Schlagworte:	Ablation Algorithms Annotations Artificial neural networks Encoders-Decoders Free form Humans Image analysis Image degradation Image processing Image segmentation Language Machine learning Medical imaging Natural Language Processing Neural networks Neural Networks, Computer Performance degradation Performance evaluation Physicians Pneumothorax Pneumothorax - diagnostic imaging Radiography, Thoracic Radiology Segmentation
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	Radiology narrative reports often describe characteristics of a patient's disease, including its location, size, and shape. Motivated by the recent success of multimodal learning, we hypothesized that this descriptive text could guide medical image analysis algorithms. We proposed a novel vision-language model, ConTEXTual Net, for the task of pneumothorax segmentation on chest radiographs. ConTEXTual Net extracts language features from physician-generated free-form radiology reports using a pre-trained language model. We then introduced cross-attention between the language features and the intermediate embeddings of an encoder-decoder convolutional neural network to enable language guidance for image analysis. ConTEXTual Net was trained on the CANDID-PTX dataset consisting of 3196 positive cases of pneumothorax with segmentation annotations from 6 different physicians as well as clinical radiology reports. Using cross-validation, ConTEXTual Net achieved a Dice score of 0.716±0.016, which was similar to the degree of inter-reader variability (0.712±0.044) computed on a subset of the data. It outperformed vision-only models (Swin UNETR: 0.670±0.015, ResNet50 U-Net: 0.677±0.015, GLoRIA: 0.686±0.014, and nnUNet 0.694±0.016) and a competing vision-language model (LAVT: 0.706±0.009). Ablation studies confirmed that it was the text information that led to the performance gains. Additionally, we show that certain augmentation methods degraded ConTEXTual Net's segmentation performance by breaking the image-text concordance. We also evaluated the effects of using different language models and activation functions in the cross-attention module, highlighting the efficacy of our chosen architectural design.
ISSN:	2948-2933 0897-1889 2948-2925 2948-2933 1618-727X
DOI:	10.1007/s10278-024-01051-8