[1] 3
The R console window is the left (or lower-left) window in RStudio. The R console uses a “read, eval, print” loop. This is sometimes called a REPL.
3 is the answer
Ignore the [1] for now.
R performs operations (called functions) on data and values
These can be composed arbitrarily
?function_name gives you information about what the function doesDiscuss in your groups.
Solutions to a polynomial equation \(ax^2 + bx + c = 0\) are given by \[x = \frac{-b \pm \sqrt{b^2 - 4ac}}{2a}\]
Figure out how to use R functions and operations for square roots, exponentiation, and multiplication to calculate x given a=3, b=14, c=-5 (no LLMs).
tidyverse package has a function called read_csv() that lets you read csv (comma-separated values) files into R.Error in read_csv("https://raw.githubusercontent.com/alejandroschuler/r4ds-courses/mph-nhanes/data/nhanes/nhanes_subset.csv"): could not find function "read_csv"
tidyverse packageread_csv() requires you to tell it where to find the file you want to read in
"C:\Users\me\Desktop\myfile.csv""/Users/me/Desktop/myfile.csv""http://www.mywebsite.com/myfile.csv"nhanes is now a dataset loaded into R. To look at it, just type# A tibble: 59 x 11
id age sex race diabetes glucose hba1c bmi waist_cm bp_sys_1
<chr> <dbl> <chr> <chr> <chr> <dbl> <dbl> <dbl> <dbl> <dbl>
1 NHANES-1001~ 72 male Mexi~ Yes 124 6.8 31.5 107. 124
2 NHANES-1001~ 80 fema~ White Yes 95 6.7 26 87 180
3 NHANES-1002~ 66 male White Yes 171 7.9 26.1 105. 120
4 NHANES-1002~ 65 fema~ Black Yes 157 7.4 34.4 102. 102
5 NHANES-1004~ 31 male White No 98 4.8 23.2 85.3 112
6 NHANES-1008~ 80 male Mexi~ Yes 174 7.7 32.1 120. 156
7 NHANES-1011~ 49 fema~ Black No 107 5.7 39.3 112. 128
8 NHANES-1012~ 65 fema~ Black Yes 150 8.1 30.8 114. 128
9 NHANES-1013~ 80 fema~ White Yes 165 8.7 30 118. 156
10 NHANES-1014~ 53 male Asian No 104 5.6 28.6 99.1 126
# i 49 more rows
# i 1 more variable: cholesterol <dbl>
This is a data frame, one of the most powerful features in R (a “tibble” is a kind of data frame). - Similar to an Excel spreadsheet. - One row ~ one instance of some (real-world) object. - One column ~ one variable, containing the values for the corresponding instances. - All the values in one column should be of the same type (a number, a category, text, etc.), but different columns can be of different types.
# A tibble: 59 x 11
id age sex race diabetes glucose hba1c bmi waist_cm bp_sys_1
<chr> <dbl> <chr> <chr> <chr> <dbl> <dbl> <dbl> <dbl> <dbl>
1 NHANES-1001~ 72 male Mexi~ Yes 124 6.8 31.5 107. 124
2 NHANES-1001~ 80 fema~ White Yes 95 6.7 26 87 180
3 NHANES-1002~ 66 male White Yes 171 7.9 26.1 105. 120
4 NHANES-1002~ 65 fema~ Black Yes 157 7.4 34.4 102. 102
5 NHANES-1004~ 31 male White No 98 4.8 23.2 85.3 112
6 NHANES-1008~ 80 male Mexi~ Yes 174 7.7 32.1 120. 156
7 NHANES-1011~ 49 fema~ Black No 107 5.7 39.3 112. 128
8 NHANES-1012~ 65 fema~ Black Yes 150 8.1 30.8 114. 128
9 NHANES-1013~ 80 fema~ White Yes 165 8.7 30 118. 156
10 NHANES-1014~ 53 male Asian No 104 5.6 28.6 99.1 126
# i 49 more rows
# i 1 more variable: cholesterol <dbl>
This is a subset of NHANES, the CDC’s National Health and Nutrition Examination Survey. NHANES sends a mobile clinic around the country, interviews a nationally representative sample of Americans, and gives each of them a physical exam and a blood draw.
Let’s say we’re curious about the relationship between two lab values, glucose and hba1c.
ggplot(dataset) says “start a chart with this dataset”+ geom_point(...) says “put points on this chart”aes(x=x_values y=y_values) says “map the values in the column x_values to the x-axis, and map the values in the column y_values to the y-axis” (aes is short for aesthetic)ggplot is short for “grammar of graphics plot”
ggplot() and geom_point() are functions imported from the ggplot2 package, which is one of the “sub-packages” of the tidyverse package we loaded earlierMake a scatterplot of diabetes vs hba1c (another lab value in the dataset). The result should look like this:
Let’s say we’re curious about the relationship between glucose and hba1c.
Can you recreate this plot?
What will this do? Why?
?geom_point to see what aesthetics are expected or allowedggplot function directly instead of to each geom individuallyUse google or other resources to figure out how to receate this plot in R:
ggplot(nhanes) +
...
facet_wrap is good for faceting according to unordered categoriesfacet_grid is better for ordered categories, and can be used with two variablesUse ggplot to investigate the relationship between these measurements and diabetes using any combination of any kinds of plots that you like. Which measurements are most associated with diabetes? Does this vary by race, age, or assigned sex at birth?
For some plots it may be helpful to reformat your data using this code (we’ll learn how to do this on day 4):
# A tibble: 354 x 7
id age sex race diabetes measure value
<chr> <dbl> <chr> <chr> <chr> <chr> <dbl>
1 NHANES-100129 72 male Mexican American Yes glucose 124
2 NHANES-100129 72 male Mexican American Yes hba1c 6.8
3 NHANES-100129 72 male Mexican American Yes bmi 31.5
4 NHANES-100129 72 male Mexican American Yes waist_cm 107.
5 NHANES-100129 72 male Mexican American Yes bp_sys_1 124
6 NHANES-100129 72 male Mexican American Yes cholesterol 170
7 NHANES-100188 80 female White Yes glucose 95
8 NHANES-100188 80 female White Yes hba1c 6.7
9 NHANES-100188 80 female White Yes bmi 26
10 NHANES-100188 80 female White Yes waist_cm 87
# i 344 more rows