Coding with the machines: machine-assisted coding of rare event data

While machine coding of data has dramatically advanced in recent years, the literature raises significant concerns about validation of LLM classification showing, for example, that reliability varies greatly by prompt and temperature tuning, across subject areas and tasks-especially in "zero-sh...

Ausführliche Beschreibung

Gespeichert in:
Bibliographische Detailangaben
Veröffentlicht in:PNAS nexus 2024-05, Vol.3 (5), p.pgae165-pgae165
Hauptverfasser: Overos, Henry David, Hlatky, Roman, Pathak, Ojashwi, Goers, Harriet, Gouws-Dewar, Jordan, Smith, Katy, Chew, Keith Padraic, Birnir, Jóhanna K, Liu, Amy H
Format: Artikel
Sprache:eng
Schlagworte:
Online-Zugang:Volltext
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
container_end_page pgae165
container_issue 5
container_start_page pgae165
container_title PNAS nexus
container_volume 3
creator Overos, Henry David
Hlatky, Roman
Pathak, Ojashwi
Goers, Harriet
Gouws-Dewar, Jordan
Smith, Katy
Chew, Keith Padraic
Birnir, Jóhanna K
Liu, Amy H
description While machine coding of data has dramatically advanced in recent years, the literature raises significant concerns about validation of LLM classification showing, for example, that reliability varies greatly by prompt and temperature tuning, across subject areas and tasks-especially in "zero-shot" applications. This paper contributes to the discussion of validation in several different ways. To test the relative performance of supervised and semi-supervised algorithms when coding political data, we compare three models' performances to each other over multiple iterations for each model and to trained expert coding of data. We also examine changes in performance resulting from prompt engineering and pre-processing of source data. To ameliorate concerns regarding LLM's pre-training on test data, we assess performance by updating an existing dataset beyond what is publicly available. Overall, we find that only GPT-4 approaches trained expert coders when coding contexts familiar to human coders and codes more consistently across contexts. We conclude by discussing some benefits and drawbacks of machine coding moving forward.
doi_str_mv 10.1093/pnasnexus/pgae165
format Article
fullrecord <record><control><sourceid>gale_pubme</sourceid><recordid>TN_cdi_pubmedcentral_primary_oai_pubmedcentral_nih_gov_11102067</recordid><sourceformat>XML</sourceformat><sourcesystem>PC</sourcesystem><galeid>A800097416</galeid><sourcerecordid>A800097416</sourcerecordid><originalsourceid>FETCH-LOGICAL-c419t-1a3990bdf53de097448122eb9ab846c090711493a0cd1f205aaec4579f403363</originalsourceid><addsrcrecordid>eNptUU1rGzEQFaWlCWl-QC9loZdeNpmRViurlxLcTwj0krsYa2dtlV3JXa2T9t9Xxo5JoOgwg-a9N096QrxFuEKw6nobKUf-s8vX2zUxtvqFOJdGy7rVjXz5pD8Tlzn_AgBpDGKjX4sztTCtNqjPxedl6kJcVw9h3lTzhquR_CZEzh8fu5pyDnnmrvIHaOqriSau-J7jXHU00xvxqqch8-WxXoi7r1_ult_r25_ffixvbmvfoJ1rJGUtrLpeq47BmqZZoJS8srRaNK0HC3t_VhH4DnsJmoh9o43tG1CqVRfi00F2u1uN3PmyfqLBbacw0vTXJQru-SSGjVune4eIIKE1ReHDUWFKv3ecZzeG7HkYKHLaZadAGzASAQv0_QG6poFdiH0qkn4PdzeL8pnFPu4tXf0HVU7HY_Apch_K_TMCHgh-SjlP3J_sI7h9ru6UqzvmWjjvnr77xHhMUf0DKpyghA</addsrcrecordid><sourcetype>Open Access Repository</sourcetype><iscdi>true</iscdi><recordtype>article</recordtype><pqid>3057072101</pqid></control><display><type>article</type><title>Coding with the machines: machine-assisted coding of rare event data</title><source>DOAJ Directory of Open Access Journals</source><source>Oxford Journals Open Access Collection</source><source>EZB-FREE-00999 freely available EZB journals</source><source>PubMed Central</source><creator>Overos, Henry David ; Hlatky, Roman ; Pathak, Ojashwi ; Goers, Harriet ; Gouws-Dewar, Jordan ; Smith, Katy ; Chew, Keith Padraic ; Birnir, Jóhanna K ; Liu, Amy H</creator><contributor>Ognyanova, Katherine</contributor><creatorcontrib>Overos, Henry David ; Hlatky, Roman ; Pathak, Ojashwi ; Goers, Harriet ; Gouws-Dewar, Jordan ; Smith, Katy ; Chew, Keith Padraic ; Birnir, Jóhanna K ; Liu, Amy H ; Ognyanova, Katherine</creatorcontrib><description>While machine coding of data has dramatically advanced in recent years, the literature raises significant concerns about validation of LLM classification showing, for example, that reliability varies greatly by prompt and temperature tuning, across subject areas and tasks-especially in "zero-shot" applications. This paper contributes to the discussion of validation in several different ways. To test the relative performance of supervised and semi-supervised algorithms when coding political data, we compare three models' performances to each other over multiple iterations for each model and to trained expert coding of data. We also examine changes in performance resulting from prompt engineering and pre-processing of source data. To ameliorate concerns regarding LLM's pre-training on test data, we assess performance by updating an existing dataset beyond what is publicly available. Overall, we find that only GPT-4 approaches trained expert coders when coding contexts familiar to human coders and codes more consistently across contexts. We conclude by discussing some benefits and drawbacks of machine coding moving forward.</description><identifier>ISSN: 2752-6542</identifier><identifier>EISSN: 2752-6542</identifier><identifier>DOI: 10.1093/pnasnexus/pgae165</identifier><identifier>PMID: 38765715</identifier><language>eng</language><publisher>England: Oxford University Press</publisher><subject>Electronic data processing ; Machine learning ; Methods ; Social and Political Sciences</subject><ispartof>PNAS nexus, 2024-05, Vol.3 (5), p.pgae165-pgae165</ispartof><rights>The Author(s) 2024. Published by Oxford University Press on behalf of National Academy of Sciences.</rights><rights>COPYRIGHT 2024 Oxford University Press</rights><rights>The Author(s) 2024. Published by Oxford University Press on behalf of National Academy of Sciences. 2024</rights><lds50>peer_reviewed</lds50><oa>free_for_read</oa><woscitedreferencessubscribed>false</woscitedreferencessubscribed><cites>FETCH-LOGICAL-c419t-1a3990bdf53de097448122eb9ab846c090711493a0cd1f205aaec4579f403363</cites><orcidid>0000-0001-6010-4707 ; 0000-0001-8378-2877 ; 0009-0007-7653-6214 ; 0000-0001-9804-7752 ; 0009-0006-3251-055X ; 0000-0003-1789-0268 ; 0000-0003-1261-8812 ; 0000-0001-5380-2849</orcidid></display><links><openurl>$$Topenurl_article</openurl><openurlfulltext>$$Topenurlfull_article</openurlfulltext><thumbnail>$$Tsyndetics_thumb_exl</thumbnail><linktopdf>$$Uhttps://www.ncbi.nlm.nih.gov/pmc/articles/PMC11102067/pdf/$$EPDF$$P50$$Gpubmedcentral$$Hfree_for_read</linktopdf><linktohtml>$$Uhttps://www.ncbi.nlm.nih.gov/pmc/articles/PMC11102067/$$EHTML$$P50$$Gpubmedcentral$$Hfree_for_read</linktohtml><link.rule.ids>230,314,727,780,784,864,885,27924,27925,53791,53793</link.rule.ids><backlink>$$Uhttps://www.ncbi.nlm.nih.gov/pubmed/38765715$$D View this record in MEDLINE/PubMed$$Hfree_for_read</backlink></links><search><contributor>Ognyanova, Katherine</contributor><creatorcontrib>Overos, Henry David</creatorcontrib><creatorcontrib>Hlatky, Roman</creatorcontrib><creatorcontrib>Pathak, Ojashwi</creatorcontrib><creatorcontrib>Goers, Harriet</creatorcontrib><creatorcontrib>Gouws-Dewar, Jordan</creatorcontrib><creatorcontrib>Smith, Katy</creatorcontrib><creatorcontrib>Chew, Keith Padraic</creatorcontrib><creatorcontrib>Birnir, Jóhanna K</creatorcontrib><creatorcontrib>Liu, Amy H</creatorcontrib><title>Coding with the machines: machine-assisted coding of rare event data</title><title>PNAS nexus</title><addtitle>PNAS Nexus</addtitle><description>While machine coding of data has dramatically advanced in recent years, the literature raises significant concerns about validation of LLM classification showing, for example, that reliability varies greatly by prompt and temperature tuning, across subject areas and tasks-especially in "zero-shot" applications. This paper contributes to the discussion of validation in several different ways. To test the relative performance of supervised and semi-supervised algorithms when coding political data, we compare three models' performances to each other over multiple iterations for each model and to trained expert coding of data. We also examine changes in performance resulting from prompt engineering and pre-processing of source data. To ameliorate concerns regarding LLM's pre-training on test data, we assess performance by updating an existing dataset beyond what is publicly available. Overall, we find that only GPT-4 approaches trained expert coders when coding contexts familiar to human coders and codes more consistently across contexts. We conclude by discussing some benefits and drawbacks of machine coding moving forward.</description><subject>Electronic data processing</subject><subject>Machine learning</subject><subject>Methods</subject><subject>Social and Political Sciences</subject><issn>2752-6542</issn><issn>2752-6542</issn><fulltext>true</fulltext><rsrctype>article</rsrctype><creationdate>2024</creationdate><recordtype>article</recordtype><recordid>eNptUU1rGzEQFaWlCWl-QC9loZdeNpmRViurlxLcTwj0krsYa2dtlV3JXa2T9t9Xxo5JoOgwg-a9N096QrxFuEKw6nobKUf-s8vX2zUxtvqFOJdGy7rVjXz5pD8Tlzn_AgBpDGKjX4sztTCtNqjPxedl6kJcVw9h3lTzhquR_CZEzh8fu5pyDnnmrvIHaOqriSau-J7jXHU00xvxqqch8-WxXoi7r1_ult_r25_ffixvbmvfoJ1rJGUtrLpeq47BmqZZoJS8srRaNK0HC3t_VhH4DnsJmoh9o43tG1CqVRfi00F2u1uN3PmyfqLBbacw0vTXJQru-SSGjVune4eIIKE1ReHDUWFKv3ecZzeG7HkYKHLaZadAGzASAQv0_QG6poFdiH0qkn4PdzeL8pnFPu4tXf0HVU7HY_Apch_K_TMCHgh-SjlP3J_sI7h9ru6UqzvmWjjvnr77xHhMUf0DKpyghA</recordid><startdate>20240501</startdate><enddate>20240501</enddate><creator>Overos, Henry David</creator><creator>Hlatky, Roman</creator><creator>Pathak, Ojashwi</creator><creator>Goers, Harriet</creator><creator>Gouws-Dewar, Jordan</creator><creator>Smith, Katy</creator><creator>Chew, Keith Padraic</creator><creator>Birnir, Jóhanna K</creator><creator>Liu, Amy H</creator><general>Oxford University Press</general><scope>NPM</scope><scope>AAYXX</scope><scope>CITATION</scope><scope>7X8</scope><scope>5PM</scope><orcidid>https://orcid.org/0000-0001-6010-4707</orcidid><orcidid>https://orcid.org/0000-0001-8378-2877</orcidid><orcidid>https://orcid.org/0009-0007-7653-6214</orcidid><orcidid>https://orcid.org/0000-0001-9804-7752</orcidid><orcidid>https://orcid.org/0009-0006-3251-055X</orcidid><orcidid>https://orcid.org/0000-0003-1789-0268</orcidid><orcidid>https://orcid.org/0000-0003-1261-8812</orcidid><orcidid>https://orcid.org/0000-0001-5380-2849</orcidid></search><sort><creationdate>20240501</creationdate><title>Coding with the machines: machine-assisted coding of rare event data</title><author>Overos, Henry David ; Hlatky, Roman ; Pathak, Ojashwi ; Goers, Harriet ; Gouws-Dewar, Jordan ; Smith, Katy ; Chew, Keith Padraic ; Birnir, Jóhanna K ; Liu, Amy H</author></sort><facets><frbrtype>5</frbrtype><frbrgroupid>cdi_FETCH-LOGICAL-c419t-1a3990bdf53de097448122eb9ab846c090711493a0cd1f205aaec4579f403363</frbrgroupid><rsrctype>articles</rsrctype><prefilter>articles</prefilter><language>eng</language><creationdate>2024</creationdate><topic>Electronic data processing</topic><topic>Machine learning</topic><topic>Methods</topic><topic>Social and Political Sciences</topic><toplevel>peer_reviewed</toplevel><toplevel>online_resources</toplevel><creatorcontrib>Overos, Henry David</creatorcontrib><creatorcontrib>Hlatky, Roman</creatorcontrib><creatorcontrib>Pathak, Ojashwi</creatorcontrib><creatorcontrib>Goers, Harriet</creatorcontrib><creatorcontrib>Gouws-Dewar, Jordan</creatorcontrib><creatorcontrib>Smith, Katy</creatorcontrib><creatorcontrib>Chew, Keith Padraic</creatorcontrib><creatorcontrib>Birnir, Jóhanna K</creatorcontrib><creatorcontrib>Liu, Amy H</creatorcontrib><collection>PubMed</collection><collection>CrossRef</collection><collection>MEDLINE - Academic</collection><collection>PubMed Central (Full Participant titles)</collection><jtitle>PNAS nexus</jtitle></facets><delivery><delcategory>Remote Search Resource</delcategory><fulltext>fulltext</fulltext></delivery><addata><au>Overos, Henry David</au><au>Hlatky, Roman</au><au>Pathak, Ojashwi</au><au>Goers, Harriet</au><au>Gouws-Dewar, Jordan</au><au>Smith, Katy</au><au>Chew, Keith Padraic</au><au>Birnir, Jóhanna K</au><au>Liu, Amy H</au><au>Ognyanova, Katherine</au><format>journal</format><genre>article</genre><ristype>JOUR</ristype><atitle>Coding with the machines: machine-assisted coding of rare event data</atitle><jtitle>PNAS nexus</jtitle><addtitle>PNAS Nexus</addtitle><date>2024-05-01</date><risdate>2024</risdate><volume>3</volume><issue>5</issue><spage>pgae165</spage><epage>pgae165</epage><pages>pgae165-pgae165</pages><issn>2752-6542</issn><eissn>2752-6542</eissn><abstract>While machine coding of data has dramatically advanced in recent years, the literature raises significant concerns about validation of LLM classification showing, for example, that reliability varies greatly by prompt and temperature tuning, across subject areas and tasks-especially in "zero-shot" applications. This paper contributes to the discussion of validation in several different ways. To test the relative performance of supervised and semi-supervised algorithms when coding political data, we compare three models' performances to each other over multiple iterations for each model and to trained expert coding of data. We also examine changes in performance resulting from prompt engineering and pre-processing of source data. To ameliorate concerns regarding LLM's pre-training on test data, we assess performance by updating an existing dataset beyond what is publicly available. Overall, we find that only GPT-4 approaches trained expert coders when coding contexts familiar to human coders and codes more consistently across contexts. We conclude by discussing some benefits and drawbacks of machine coding moving forward.</abstract><cop>England</cop><pub>Oxford University Press</pub><pmid>38765715</pmid><doi>10.1093/pnasnexus/pgae165</doi><orcidid>https://orcid.org/0000-0001-6010-4707</orcidid><orcidid>https://orcid.org/0000-0001-8378-2877</orcidid><orcidid>https://orcid.org/0009-0007-7653-6214</orcidid><orcidid>https://orcid.org/0000-0001-9804-7752</orcidid><orcidid>https://orcid.org/0009-0006-3251-055X</orcidid><orcidid>https://orcid.org/0000-0003-1789-0268</orcidid><orcidid>https://orcid.org/0000-0003-1261-8812</orcidid><orcidid>https://orcid.org/0000-0001-5380-2849</orcidid><oa>free_for_read</oa></addata></record>
fulltext fulltext
identifier ISSN: 2752-6542
ispartof PNAS nexus, 2024-05, Vol.3 (5), p.pgae165-pgae165
issn 2752-6542
2752-6542
language eng
recordid cdi_pubmedcentral_primary_oai_pubmedcentral_nih_gov_11102067
source DOAJ Directory of Open Access Journals; Oxford Journals Open Access Collection; EZB-FREE-00999 freely available EZB journals; PubMed Central
subjects Electronic data processing
Machine learning
Methods
Social and Political Sciences
title Coding with the machines: machine-assisted coding of rare event data
url https://sfx.bib-bvb.de/sfx_tum?ctx_ver=Z39.88-2004&ctx_enc=info:ofi/enc:UTF-8&ctx_tim=2024-12-19T16%3A19%3A16IST&url_ver=Z39.88-2004&url_ctx_fmt=infofi/fmt:kev:mtx:ctx&rfr_id=info:sid/primo.exlibrisgroup.com:primo3-Article-gale_pubme&rft_val_fmt=info:ofi/fmt:kev:mtx:journal&rft.genre=article&rft.atitle=Coding%20with%20the%20machines:%20machine-assisted%20coding%20of%20rare%20event%20data&rft.jtitle=PNAS%20nexus&rft.au=Overos,%20Henry%20David&rft.date=2024-05-01&rft.volume=3&rft.issue=5&rft.spage=pgae165&rft.epage=pgae165&rft.pages=pgae165-pgae165&rft.issn=2752-6542&rft.eissn=2752-6542&rft_id=info:doi/10.1093/pnasnexus/pgae165&rft_dat=%3Cgale_pubme%3EA800097416%3C/gale_pubme%3E%3Curl%3E%3C/url%3E&disable_directlink=true&sfx.directlink=off&sfx.report_link=0&rft_id=info:oai/&rft_pqid=3057072101&rft_id=info:pmid/38765715&rft_galeid=A800097416&rfr_iscdi=true