R for Health Data Science
Book information
Description
In this age of information, the manipulation, analysis, and interpretation of data have become a fundamental part of professional life; nowhere more so than in the delivery of healthcare. From the understanding of disease and the development of new treatments, to the diagnosis and management of individual patients, the use of data and technology is now an integral part of the business of healthcare. Those working in healthcare interact daily with data, often without realising it. The conversion of this avalanche of information to useful knowledge is essential for high-quality patient care. "R for Health Data Science" includes everything a healthcare professional needs to go from R novice to R guru. By the end of this book, you will be taking a sophisticated approach to health data science with beautiful visualisations, elegant tables, and nuanced analyses. Preface About the Authors I Data wrangling and visualisation Why we love R Help, what's a script? What is RStudio? Getting started Getting help Work in a Project Restart R regularly Notation throughout this book R basics Reading data into R Import Dataset interface Reading in the Global Burden of Disease example dataset Variable types and why we care Numeric variables (continuous) Character variables Factor variables (categorical) Date/time variables Objects and functions data frame/tibble Naming objects Function and its arguments Working with objects % Using . to direct the pipe Operators for filtering data Worked examples The combine function: c() Missing values (NAs) and filters Creating new columns - mutate() Worked example/exercise Conditional calculations - if_else() Create labels - paste() Joining multiple datasets Further notes about joins Summarising data Get the data Plot the data Aggregating: group_by(), summarise() Add new columns: mutate() Percentages formatting: percent() summarise() vs mutate() Common arithmetic functions - sum(), mean(), median(), etc. select() columns Reshaping data - long vs wide format Pivot values from rows into columns (wider) Pivot values from columns to rows (longer) separate() a column into multiple columns arrange() rows Factor levels Exercises Exercise - pivot_wider() Exercise - group_by(), summarise() Exercise - full_join(), percent() Exercise - mutate(), summarise() Exercise - filter(), summarise(), pivot_wider() Different types of plots Get the data Anatomy of ggplot explained Set your theme - grey vs white Scatter plots/bubble plots Line plots/time series plots Exercise Bar plots Summarised data Countable data colour vs fill Proportions Exercise Histograms Box plots Multiple geoms, multiple aes() Worked example - three geoms together All other types of plots Solutions Extra: Advanced examples Fine tuning plots Get the data Scales Logarithmic Expand limits Zoom in Exercise Axis ticks Colours Using the Brewer palettes: Legend title Choosing colours manually Titles and labels Annotation Annotation with a superscript and a variable Overall look - theme() Text size Legend position Saving your plot II Data analysis Working with continuous outcome variables Continuous data The Question Get and check the data Plot the data Histogram Quantile-quantile (Q-Q) plot Boxplot Compare the means of two groups t-test Two-sample t-tests Paired t-tests What if I run the wrong test? Compare the mean of one group: one sample t-tests Interchangeability of t-tests Compare the means of more than two groups Plot the data ANOVA Assumptions Multiple testing Pairwise testing and multiple comparisons Non-parametric tests Transforming data Non-parametric test for comparing two groups Non-parametric test for comparing more than two groups Finalfit approach Conclusions Exercises Exercise Exercise Exercise Exercise Solutions Linear regression Regression The Question (1) Fitting a regression line When the line fits well The fitted line and the linear equation Effect modification R-squared and model fit Confounding Summary Fitting simple models The Question (2) Get the data Check the data Plot the data Simple linear regression Multivariable linear regression Check assumptions Fitting more complex models The Question (3) Model fitting principles AIC Get the data Check the data Plot the data Linear regression with finalfit Summary Exercises Exercise Exercise Exercise Exercise Solutions Working with categorical outcome variables Factors The Question Get the data Check the data Recode the data Should I convert a continuous variable to a categorical variable? Equal intervals vs quantiles Plot the data Group factor levels together - fct_collapse() Change the order of values within a factor - fct_relevel() Summarising factors with finalfit Pearson's chi-squared and Fisher's exact tests Base R Fisher's exact test Chi-squared / Fisher's exact test using finalfit Exercises Exercise Exercise Exercise Logistic regression Generalised linear modelling Binary logistic regression The Question (1) Odds and probabilities Odds ratios Fitting a regression line The fitted line and the logistic regression equation Effect modification and confounding Data preparation and exploratory analysis The Question (2) Get the data Check the data Recode the data Plot the data Tabulate data Model assumptions Linearity of continuous variables to the response Multicollinearity Fitting logistic regression models in base R Modelling strategy for binary outcomes Fitting logistic regression models with finalfit Criterion-based model fitting Model fitting Odds ratio plot Correlated groups of observations Simulate data Plot the data Mixed effects models in base R Exercises Exercise Exercise Exercise Exercise Solutions Time-to-event data and survival The Question Get and check the data Death status Time and censoring Recode the data Kaplan Meier survival estimator KM analysis for whole cohort Model Life table Kaplan Meier plot Cox proportional hazards regression coxph() finalfit() Reduced model Testing for proportional hazards Stratified models Correlated groups of observations Hazard ratio plot Competing risks regression Summary Dates in R Converting dates to survival time Exercises Exercise Exercise Solutions III Workflow The problem of missing data Identification of missing data Missing completely at random (MCAR) Missing at random (MAR) Missing not at random (MNAR) Ensure your data are coded correctly: ff_glimpse() The Question Identify missing values in each variable: missing_plot() Look for patterns of missingness: missing_pattern() Including missing data in demographics tables Check for associations between missing and observed data For those who like an omnibus test Handling missing data: MCAR Common solution: row-wise deletion Other considerations Handling missing data: MAR Common solution: Multivariate Imputation by Chained Equations (mice) Handling missing data: MNAR Summary Notebooks and Markdown What is a Notebook? What is Markdown? What is the difference between a Notebook and an R Markdown file? Notebook vs HTML vs PDF vs Word The anatomy of a Notebook / R Markdown file YAML header R code chunks Setting default chunk options Setting default figure options Markdown elements Interface and outputting Running code and chunks, knitting File structure and workflow Why go to all this bother? Exporting and reporting Which format should I use? Working in a .R file Demographics table Logistic regression table Odds ratio plot MS Word via knitr/R Markdown Figure quality in Word output Create Word template file PDF via knitr/R Markdown Working in a .Rmd file Moving between formats Summary Version control Setup Git on RStudio and associate with GitHub Create an SSH RSA key and add to your GitHub account Create a project in RStudio and commit a file Create a new repository on GitHub and link to RStudio project Clone an existing GitHub project to new RStudio project Summary Encryption Safe practice encryptr package Get the package Get the data Generate private/public keys Encrypt columns of data Decrypt specific information only Using a lookup table Encrypting a file Decrypting a file Ciphertexts are not matchable Providing a public key Use cases Blinding in trials Re-contacting participants Long-term follow-up of participants Summary Appendix Bibliography Index
Similar books
R for Health Data Science
2020 · EPUB
MySQL® Notes for Professionals book
2018 · PDF
MrExcel 2022: Boosting Excel
2022 · PDF
MrExcel 2022: Boosting Excel
2022 · PDF
Session C11: Ancient Cultural Landscapes in South Europe – their Ecological Setting and Evolution, Session C22: Gardeners from South America, Session S04: Agro-Pastoralism and Early Metallurgy Sessions, Session WS29: The Idea of Enclosure in Recent Iberian Prehistory, Session C88: Rhytmes et causalites des dynamiques de l'anthropisation en Europe entre 6500 ET 500 BC: Hypotheses socio-culturelles et/ou climatiques: Proceedings of the XV UISPP World Congress (Lisbon 4-9 September 2006) / Actes du XV Congrès Mondial (Lisbonne 4-9 Septembre 2006) Vol.36
2010 · PDF
THE BRITISH ARMY IN INDIA: ITS PRESERVATION BY AN APPROPRIATE CLOTHING, HOUSING, LOCATING, RECREATIVE EMPLOYMENT, AND HOPEFUL ENCOURAGEMENT OF THE TROOPS. with AN APPENDIX ON INDIA : THE CLIMATE OP ITS HILLS ; THE DEVELOPMENT OF ITS RESODRCBS, INDUSTRY, AND ARTS ; THE ADMINISTRATION OF JUSTICE ; THE BLACK ACT ; THE PROGRESS OF CHRISTIANITY ; THE TRAFFIC IN OPIUM ; THE VALUE OF INDIA ; PERMANENT CAUSES OF DISAFFECTION, AND OF THE RECENT REBELLION ; THE TRADITIONARY POLICY; MISGOVERNMENT BY NATIVE RULERS ; ANNEXATIONS OF THEIR TERRITORY, ETC.
1858 · PDF
Idries Shah 27 Books Collection : A Perfumed Scorpion, A Veiled Gazelle, Caravan of Dreams, Darkest England, Destination Mecca, Evenings with Idries Shah, Knowing How to Know, Learning How to Learn, Letters and Lectures of Idries Shah, Neglected aspects of Sufi study, Observations, Oriental Magic, Reflections, Seeker after Truth, Special Illumination, Special Problems in the study of Sufi ideas, Sufi thought and action, Tales of the Dervishes, The Dermis Probe, The Elephant in the Dark, The Englishman Handbook, Idries Shah Antology, The Magic Monastery, The natives are restless, wisdom of the Idiots PDF.
2022 · PDF
The travels of Capts. Lewis and Clarke from St. Louis, by way of the Missouri and Columbia rivers, to the Pacific ocean; performed in the years 1804, 1805 & 1806, by order of the government of the United States. Containing delineations of the manners, customs, religion, &c. of the Indians, comp. from various authentic sources, and original documents, and a summary of the Statistical view of the Indian nations, from the official communication of Meriwether Lewis. Illustrated with a map of the country, inhabited by the western tribes of Indians
1809 · PDF