Week 4 Tutorial: Functions in R

POP77001 Computer Programming for Social Scientists

Naming Functions

  • Use lower_case_with_underscores style for R objects.
  • Give short, but descriptive names.
  • Try to use verbs in function names:
summarise()
get_coefs()

rather than:

summary()
coefs()
  • An added benefit is in distinguishing user-defined functions from built-it (e.g. summary(), coef())

Code Layout

  • Limit all lines to a maximum of 79 characters.
  • If the function name and definition are too long, break it up:
    • across function-indented lines:
long_function_name <- function(a = "a long argument",
                               b = "another argument",
                               c = "another long argument") {
  # As usual code is indented by two spaces.
}
  • across double-indented lines:
long_function_name <- function(
    a = "a long argument",
    b = "another argument",
    c = "another long argument") {
  # As usual code is indented by two spaces.
}
  • When it fits, function-indented style should be preferred to double-indented.

Exercise 1: Function Definition

  • Study the code for calculating a t-test below:
t_test <- function(x, y) {
  # Calculate the means of the two samples
  mean_x <- mean(x)
  mean_y <- mean(y)
  
  # Calculate the variances of the two samples
  var_x <- var(x)
  var_y <- var(y)
  
  # Calculate the sample sizes
  n_x <- length(x)
  n_y <- length(y)
  
  # Calculate the sample standard errors
  se_x <- sqrt(var_x/n_x)
  se_y <- sqrt(var_y/n_y)
  se <- sqrt(se_x^2 + se_y^2)
  
  # Calculate the t-statistic
  t_stat <- (mean_x - mean_y) / se
  
  # Calculate the degrees of freedom
  df <- se^4/(se_x^4/(n_x - 1) + se_y^4/(n_y - 1))
  
  # Calculate the p-value
  p_val <- 2 * (1 - pt(abs(t_stat), df))
  
  # Create a result list
  res <- list(
    t_statistic = t_stat,
    degrees_of_freedom = df,
    p_value = p_val
  )
  
  return(res)
}

Exercise 1: Function Components

  • What are the components of the function t_test()?
  • Check each component using access functions.

Exercise 2: Function Call

  • Call the function on 2 random samples from the standard normal distribution (rnorm()).
  • What is the return object of this function?
  • Compare the return object of the t_test() with that of the built-in t.test().
  • Modify the function code to use implicit return.
  • Modify the function to return a named vector instead of a list.

Exercise 3: Built-in Source Code

  • As R is an open source language, you can check the inner workings of all the built-in functions.
  • Many base R function can be accessed by simply typing the function name without parentheses.
  • Alternatively, you can use View() to see the source code in a separate window.
  • Functions that are designed to work with multiple classes of objects can be accessed using getS3method().
  • Run getS3method("t.test", "default") to study the source code of the built-in t.test() function.

Exercise 4: Implementing Functions

  • Let’s re-visit the moving average example from the previous week.
  • The code snippet we worked previously with was implemented using a repeat loop.
  • Re-write the same moving average calculation using a function that internally relies on a for loop.
  • Expose the window size as a function argument, so that the user can specify the desired moving average window size.

set.seed(2026)
v <- sample(1:100, 20)
v
 [1] 93 97 38 45 91 36 48  5 31 44 58 88 19 79 10 54 18 34 12 56
# Preallocate a vector to hold the moving averages.
# As we are taking a 3-point moving average, the resulting vector
# will be of length 2 less than the original vector.
mas <- vector("double", length = length(v) - 2)

i <- 1

repeat {
  mas[i] <- mean(v[i:(i + 2)])

  i <- i + 1
  
  if (i > length(v) - 2) {
    break
  }
}
mas
 [1] 76.00000 60.00000 58.00000 57.33333 58.33333 29.66667 28.00000 26.66667
 [9] 44.33333 63.33333 55.00000 62.00000 36.00000 47.66667 27.33333 35.33333
[17] 21.33333 34.00000

Exercise 5: Functionals

  • As R is a functional language, many of iteration routines can be avoided.
  • E.g. instead of creating a loop for calculating standard deviations, we can use apply() function.
  • apply(<object_name>, 2, <function_name>) allows to calculate the desired summary statistic for each of the variables.
  • Apply this function to the matrix from the exercise above
  • Now, change 2 in the function call to 1
  • What do you see? What do the current numbers show? Does this summary make sense and why?

# When dealing with random number generation it's always a good idea to make your code replicable
# by setting the seed with set.seed(function)
set.seed(2026)
# Here we create a matrix of 30 observations of 5 variables
# where each variable is a random draw from a normal distribution with mean 0
# and standard deviation drawn from a uniform distribution between 0 and 10
mat <- mapply(
  function(x) cbind(rnorm(n = 30, mean = 0, sd = x)),
  runif(n = 5, min = 0, max = 10)
)
mat
             [,1]       [,2]        [,3]        [,4]         [,5]
 [1,] -13.6781050 -7.7087194 -2.16914035 -1.81900372   0.29705818
 [2,]   7.5797083 -0.4875531 -0.13225895 -3.04744941  -2.64631288
 [3,]   1.4249915  2.3294226 -0.31626606 -0.64116119  -7.95778099
 [4,]   3.5013766 -1.7156379  0.51118260 -3.66250553  -1.87064692
 [5,]   7.1974521  5.9590225 -1.85298111  3.90809359   2.95890958
 [6,]  -2.5641353  1.8866826 -3.69369438 -1.46252633  -5.30336834
 [7,] -21.2136266  0.5842800  0.02894001  0.37079180   7.31194516
 [8,] -14.7647722  7.1318762 -0.59166399 -2.47329693   2.08518805
 [9,]  -2.8163249  9.2386375 -0.79718249  0.95130900   1.26062841
[10,] -10.7832140  9.8841454  0.50805009  5.11239464  -3.64613906
[11,]  -6.2500272 11.1619441 -2.02015661  1.91039337  -4.57216125
[12,]  -7.5453671  8.1080630  0.49419165  5.34564427   1.32234916
[13,]   1.8015323 -4.3820201 -2.22674018  0.53783592   3.72376272
[14,]  -6.5199009  3.2787311 -2.32555196 -2.99188227  -0.31726434
[15,] -14.9551028 -3.1568328  0.12475723 -4.13497598  -4.14925033
[16,]  -1.3067312 -2.2743354  0.89705818  0.39980821   1.25072059
[17,]  -0.3995154 -1.9281338  0.72336312 -3.06521006   3.01125176
[18,]  -0.9339608 -1.8054710 -2.70262086 -1.12443647 -10.63364232
[19,]   8.0296694  5.3288677  0.01894250  0.93949092   3.95534915
[20,]  -1.0523591 -7.7443276 -1.69806156  1.86691821   2.14933375
[21,] -12.6939159 -5.0739350 -0.43339453  3.19182290   0.47367080
[22,]  -3.7734895  1.5109619 -0.18183491 -3.19960980  -8.15690788
[23,]  -5.0160724  1.2715538 -1.80765643 -2.57140115   4.04715548
[24,]   2.3650972  6.1756398  0.84511903 -2.37023116   6.62544045
[25,]   8.4355312 -2.5735107  1.10121132 -0.02472118   6.25838999
[26,]   3.9876828 -9.0699113  0.25032469  1.68025647   2.07329102
[27,] -11.7797487 -5.9644921  2.18025244 -4.18028908   0.09416032
[28,]  -5.8518655  4.2751399  0.08748227  3.54159839  -0.92129593
[29,]   8.7803765 -6.2288175 -0.85534101  0.29129352   0.06839505
[30,]  -0.7808176  4.1833822  0.46398201  2.60883474  -2.24601979

Additional Exercise

  • Re-visit the code for converting grades into marks:
convert_mark_to_grade <- function(mark) {
  if (mark >= 70) {
    grade <- "I"
  } else if (mark >= 60) {
    grade <- "II.1"
  } else if (mark >= 50) {
    grade <- "II.2"
  } else {
    grade <- "F"
  }
  grade
}
  • As discussed in the lecture, this function only takes a single mark as an input.
  • Modify this function such that if takes a vector of marks (longer than 1) as an input and returns a vector of grades.