Abstract
The aim of this chapter is to present introductory statistical concepts and illustrate diverse statistical techniques that showcase accurate and effective methods for handling challenging data. Specifically, this chapter focuses on addressing various situations, including complete or incomplete data affected by different types of measurement errors (ME) and dealing with imbalanced survey results derived from national data sources. In clinical experiments, instrumental inaccuracies, biological variation, or errors in questionnaire-based self-report data can produce significant MEs issues. Ignoring ME problems can cause bias or inconsistency of statistical decision-making schemes. We will focus on two sorts of MEs that are additive model errors and errors related to the limit of detection (LOD) due to the instrumental incapability of detecting low levels of biomarker measurements. Diverse statistical approaches have been created for analyzing data affected by MEs, including methods based on the parametric/nonparametric likelihood principles, Bayesian analysis, the single and multiple imputation techniques, and the repeated measurement design of experiment. In this framework, we first present a hybrid pooled–unpooled design as one of the strategies to evaluate data subject to MEs. This hybrid design and the classical techniques are compared to show the advantages/disadvantages of the considered methods. We note that currently the pooling technique receives more attention due to cost efficiency in testing large populations with respect to pandemic. Second, to exemplify the method on the parametric likelihood principles relevant to the LOD, in this chapter, we consider longitudinal mammary tumor development studies, where outcomes are affected by the following issues: (a) increases of missing data toward the end of the study; and (b) the presence of censored data caused by the detection limits of instrumental sensitivity. We show a test to carry out K-group comparisons based on the maximum likelihood approach. We apply the decision-making procedure using breast cancer in mice data. Finally, we pay our attention to the survey data, which may not be designed for measuring accurate effect of some factors or interventions. National-level publicly available survey data sets are feasible sources to evaluate the public health impact of the intervention and policy; we discuss procedures for accurate estimation and inference using such data sets. With respect to tools for analyzing unbalanced survey outcomes from national data resources, we examine extensive ranges of data-balancing techniques. We also discuss linearization methods by implementing influence function methods. This allows researchers to evaluate the variability for the relative risk in the complex survey settings. For an application, we present an estimation of seasonal influenza vaccine effectiveness. We show robust inferential performance across different relevant model assumptions.
| Original language | English |
|---|---|
| Title of host publication | Modern Inference Based on Health-Related Markers |
| Subtitle of host publication | Biomarkers and Statistical Decision Making |
| Publisher | Elsevier |
| Pages | 1-75 |
| Number of pages | 75 |
| ISBN (Electronic) | 9780128152478 |
| ISBN (Print) | 9780128152485 |
| DOIs | |
| State | Published - Jan 1 2024 |
Keywords
- Biomarkers
- Detection limit
- Hybrid design
- Measurement error
Fingerprint
Dive into the research topics of 'An array of statistical concepts and tools for handling challenging data'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver