Designed Sampling from Large Databases for Controlled Trials

The increasing prevalence of rich sources of data and the availability of electronic medical record databases and electronic registries opens tremendous opportunities for enhancing medical research. For example, controlled trials are ubiquitously used to investigate the effect of a medical treatment...

Ausführliche Beschreibung

Gespeichert in:
Bibliographische Detailangaben
Veröffentlicht in:arXiv.org 2015-09
Hauptverfasser: Ouyang, Liwen, Apley, Daniel W, Mehrotra, Sanjay
Format: Artikel
Sprache:eng
Schlagworte:
Online-Zugang:Volltext
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
container_end_page
container_issue
container_start_page
container_title arXiv.org
container_volume
creator Ouyang, Liwen
Apley, Daniel W
Mehrotra, Sanjay
description The increasing prevalence of rich sources of data and the availability of electronic medical record databases and electronic registries opens tremendous opportunities for enhancing medical research. For example, controlled trials are ubiquitously used to investigate the effect of a medical treatment, perhaps dependent on a set of patient covariates, and traditional approaches have relied primarily on randomized patient sampling and allocation to treatment and control group. However, when covariate data for a large cohort group of patients have already been collected and are available in a database, one can potentially design a treatment/control sample and allocation that provides far better estimates of the covariate-dependent effects of the treatment. In this paper, we develop a new approach that uses optimal design of experiments (DOE) concepts to accomplish this objective. The approach selects the patients for the treatment and control samples upfront, based on their covariate values, in a manner that optimizes the information content in the data. For the optimal sample selection, we develop simple guidelines and an optimization algorithm that provides solutions that are substantially better than random sampling. Moreover, our approach causes no sampling bias in the estimated effects, for the same reason that DOE principles do not bias estimated effects. We test our method with a simulation study based on a testbed data set containing information on the effect of statins on low-density lipoprotein (LDL) cholesterol.
format Article
fullrecord <record><control><sourceid>proquest</sourceid><recordid>TN_cdi_proquest_journals_2083134436</recordid><sourceformat>XML</sourceformat><sourcesystem>PC</sourcesystem><sourcerecordid>2083134436</sourcerecordid><originalsourceid>FETCH-proquest_journals_20831344363</originalsourceid><addsrcrecordid>eNqNyrEOgjAUQNHGxESi_EMTZ5LSV5DBDTQObrKTEgspKS2-B_9vBz_A6Qz37lgiAfKsUlIeWEo0CSFkeZFFAQm7Nobs6M2bv_S8OOtHPmCY-VPjaHijV91rMsSHgLwOfsXgXJxbtNrRie2HiEl_Htn5fmvrR7Zg-GyG1m4KG_qYOikqyEEpKOG_6wsSyzc2</addsrcrecordid><sourcetype>Aggregation Database</sourcetype><iscdi>true</iscdi><recordtype>article</recordtype><pqid>2083134436</pqid></control><display><type>article</type><title>Designed Sampling from Large Databases for Controlled Trials</title><source>Freely Accessible Journals</source><creator>Ouyang, Liwen ; Apley, Daniel W ; Mehrotra, Sanjay</creator><creatorcontrib>Ouyang, Liwen ; Apley, Daniel W ; Mehrotra, Sanjay</creatorcontrib><description>The increasing prevalence of rich sources of data and the availability of electronic medical record databases and electronic registries opens tremendous opportunities for enhancing medical research. For example, controlled trials are ubiquitously used to investigate the effect of a medical treatment, perhaps dependent on a set of patient covariates, and traditional approaches have relied primarily on randomized patient sampling and allocation to treatment and control group. However, when covariate data for a large cohort group of patients have already been collected and are available in a database, one can potentially design a treatment/control sample and allocation that provides far better estimates of the covariate-dependent effects of the treatment. In this paper, we develop a new approach that uses optimal design of experiments (DOE) concepts to accomplish this objective. The approach selects the patients for the treatment and control samples upfront, based on their covariate values, in a manner that optimizes the information content in the data. For the optimal sample selection, we develop simple guidelines and an optimization algorithm that provides solutions that are substantially better than random sampling. Moreover, our approach causes no sampling bias in the estimated effects, for the same reason that DOE principles do not bias estimated effects. We test our method with a simulation study based on a testbed data set containing information on the effect of statins on low-density lipoprotein (LDL) cholesterol.</description><identifier>EISSN: 2331-8422</identifier><language>eng</language><publisher>Ithaca: Cornell University Library, arXiv.org</publisher><subject>Algorithms ; Bias ; Cholesterol ; Computer simulation ; Design of experiments ; Electronic health records ; Health services ; Medical research ; Optimization ; Patients ; Random sampling ; Test procedures</subject><ispartof>arXiv.org, 2015-09</ispartof><rights>2015. This work is published under http://arxiv.org/licenses/nonexclusive-distrib/1.0/ (the “License”). Notwithstanding the ProQuest Terms and Conditions, you may use this content in accordance with the terms of the License.</rights><oa>free_for_read</oa><woscitedreferencessubscribed>false</woscitedreferencessubscribed></display><links><openurl>$$Topenurl_article</openurl><openurlfulltext>$$Topenurlfull_article</openurlfulltext><thumbnail>$$Tsyndetics_thumb_exl</thumbnail><link.rule.ids>777,781</link.rule.ids></links><search><creatorcontrib>Ouyang, Liwen</creatorcontrib><creatorcontrib>Apley, Daniel W</creatorcontrib><creatorcontrib>Mehrotra, Sanjay</creatorcontrib><title>Designed Sampling from Large Databases for Controlled Trials</title><title>arXiv.org</title><description>The increasing prevalence of rich sources of data and the availability of electronic medical record databases and electronic registries opens tremendous opportunities for enhancing medical research. For example, controlled trials are ubiquitously used to investigate the effect of a medical treatment, perhaps dependent on a set of patient covariates, and traditional approaches have relied primarily on randomized patient sampling and allocation to treatment and control group. However, when covariate data for a large cohort group of patients have already been collected and are available in a database, one can potentially design a treatment/control sample and allocation that provides far better estimates of the covariate-dependent effects of the treatment. In this paper, we develop a new approach that uses optimal design of experiments (DOE) concepts to accomplish this objective. The approach selects the patients for the treatment and control samples upfront, based on their covariate values, in a manner that optimizes the information content in the data. For the optimal sample selection, we develop simple guidelines and an optimization algorithm that provides solutions that are substantially better than random sampling. Moreover, our approach causes no sampling bias in the estimated effects, for the same reason that DOE principles do not bias estimated effects. We test our method with a simulation study based on a testbed data set containing information on the effect of statins on low-density lipoprotein (LDL) cholesterol.</description><subject>Algorithms</subject><subject>Bias</subject><subject>Cholesterol</subject><subject>Computer simulation</subject><subject>Design of experiments</subject><subject>Electronic health records</subject><subject>Health services</subject><subject>Medical research</subject><subject>Optimization</subject><subject>Patients</subject><subject>Random sampling</subject><subject>Test procedures</subject><issn>2331-8422</issn><fulltext>true</fulltext><rsrctype>article</rsrctype><creationdate>2015</creationdate><recordtype>article</recordtype><sourceid>ABUWG</sourceid><sourceid>AFKRA</sourceid><sourceid>AZQEC</sourceid><sourceid>BENPR</sourceid><sourceid>CCPQU</sourceid><sourceid>DWQXO</sourceid><recordid>eNqNyrEOgjAUQNHGxESi_EMTZ5LSV5DBDTQObrKTEgspKS2-B_9vBz_A6Qz37lgiAfKsUlIeWEo0CSFkeZFFAQm7Nobs6M2bv_S8OOtHPmCY-VPjaHijV91rMsSHgLwOfsXgXJxbtNrRie2HiEl_Htn5fmvrR7Zg-GyG1m4KG_qYOikqyEEpKOG_6wsSyzc2</recordid><startdate>20150922</startdate><enddate>20150922</enddate><creator>Ouyang, Liwen</creator><creator>Apley, Daniel W</creator><creator>Mehrotra, Sanjay</creator><general>Cornell University Library, arXiv.org</general><scope>8FE</scope><scope>8FG</scope><scope>ABJCF</scope><scope>ABUWG</scope><scope>AFKRA</scope><scope>AZQEC</scope><scope>BENPR</scope><scope>BGLVJ</scope><scope>CCPQU</scope><scope>DWQXO</scope><scope>HCIFZ</scope><scope>L6V</scope><scope>M7S</scope><scope>PIMPY</scope><scope>PQEST</scope><scope>PQQKQ</scope><scope>PQUKI</scope><scope>PRINS</scope><scope>PTHSS</scope></search><sort><creationdate>20150922</creationdate><title>Designed Sampling from Large Databases for Controlled Trials</title><author>Ouyang, Liwen ; Apley, Daniel W ; Mehrotra, Sanjay</author></sort><facets><frbrtype>5</frbrtype><frbrgroupid>cdi_FETCH-proquest_journals_20831344363</frbrgroupid><rsrctype>articles</rsrctype><prefilter>articles</prefilter><language>eng</language><creationdate>2015</creationdate><topic>Algorithms</topic><topic>Bias</topic><topic>Cholesterol</topic><topic>Computer simulation</topic><topic>Design of experiments</topic><topic>Electronic health records</topic><topic>Health services</topic><topic>Medical research</topic><topic>Optimization</topic><topic>Patients</topic><topic>Random sampling</topic><topic>Test procedures</topic><toplevel>online_resources</toplevel><creatorcontrib>Ouyang, Liwen</creatorcontrib><creatorcontrib>Apley, Daniel W</creatorcontrib><creatorcontrib>Mehrotra, Sanjay</creatorcontrib><collection>ProQuest SciTech Collection</collection><collection>ProQuest Technology Collection</collection><collection>Materials Science &amp; Engineering Collection</collection><collection>ProQuest Central (Alumni Edition)</collection><collection>ProQuest Central UK/Ireland</collection><collection>ProQuest Central Essentials</collection><collection>ProQuest Central</collection><collection>Technology Collection</collection><collection>ProQuest One Community College</collection><collection>ProQuest Central Korea</collection><collection>SciTech Premium Collection</collection><collection>ProQuest Engineering Collection</collection><collection>Engineering Database</collection><collection>Publicly Available Content Database</collection><collection>ProQuest One Academic Eastern Edition (DO NOT USE)</collection><collection>ProQuest One Academic</collection><collection>ProQuest One Academic UKI Edition</collection><collection>ProQuest Central China</collection><collection>Engineering Collection</collection></facets><delivery><delcategory>Remote Search Resource</delcategory><fulltext>fulltext</fulltext></delivery><addata><au>Ouyang, Liwen</au><au>Apley, Daniel W</au><au>Mehrotra, Sanjay</au><format>book</format><genre>document</genre><ristype>GEN</ristype><atitle>Designed Sampling from Large Databases for Controlled Trials</atitle><jtitle>arXiv.org</jtitle><date>2015-09-22</date><risdate>2015</risdate><eissn>2331-8422</eissn><abstract>The increasing prevalence of rich sources of data and the availability of electronic medical record databases and electronic registries opens tremendous opportunities for enhancing medical research. For example, controlled trials are ubiquitously used to investigate the effect of a medical treatment, perhaps dependent on a set of patient covariates, and traditional approaches have relied primarily on randomized patient sampling and allocation to treatment and control group. However, when covariate data for a large cohort group of patients have already been collected and are available in a database, one can potentially design a treatment/control sample and allocation that provides far better estimates of the covariate-dependent effects of the treatment. In this paper, we develop a new approach that uses optimal design of experiments (DOE) concepts to accomplish this objective. The approach selects the patients for the treatment and control samples upfront, based on their covariate values, in a manner that optimizes the information content in the data. For the optimal sample selection, we develop simple guidelines and an optimization algorithm that provides solutions that are substantially better than random sampling. Moreover, our approach causes no sampling bias in the estimated effects, for the same reason that DOE principles do not bias estimated effects. We test our method with a simulation study based on a testbed data set containing information on the effect of statins on low-density lipoprotein (LDL) cholesterol.</abstract><cop>Ithaca</cop><pub>Cornell University Library, arXiv.org</pub><oa>free_for_read</oa></addata></record>
fulltext fulltext
identifier EISSN: 2331-8422
ispartof arXiv.org, 2015-09
issn 2331-8422
language eng
recordid cdi_proquest_journals_2083134436
source Freely Accessible Journals
subjects Algorithms
Bias
Cholesterol
Computer simulation
Design of experiments
Electronic health records
Health services
Medical research
Optimization
Patients
Random sampling
Test procedures
title Designed Sampling from Large Databases for Controlled Trials
url https://sfx.bib-bvb.de/sfx_tum?ctx_ver=Z39.88-2004&ctx_enc=info:ofi/enc:UTF-8&ctx_tim=2025-01-18T11%3A30%3A45IST&url_ver=Z39.88-2004&url_ctx_fmt=infofi/fmt:kev:mtx:ctx&rfr_id=info:sid/primo.exlibrisgroup.com:primo3-Article-proquest&rft_val_fmt=info:ofi/fmt:kev:mtx:book&rft.genre=document&rft.atitle=Designed%20Sampling%20from%20Large%20Databases%20for%20Controlled%20Trials&rft.jtitle=arXiv.org&rft.au=Ouyang,%20Liwen&rft.date=2015-09-22&rft.eissn=2331-8422&rft_id=info:doi/&rft_dat=%3Cproquest%3E2083134436%3C/proquest%3E%3Curl%3E%3C/url%3E&disable_directlink=true&sfx.directlink=off&sfx.report_link=0&rft_id=info:oai/&rft_pqid=2083134436&rft_id=info:pmid/&rfr_iscdi=true