4  Tidy Time Series Basics

At its core, a time series is just a data where when it was collected is a crucial part of the story. It is a sequence of data points indexed in time and order. Regardless, a time series should still be a tidy data frame. In a tidy dataset (following Hadley Wickham’s principles), Each row is an observation, each column is a variable.

4.1 What is a tsibble

A tsibble (short for “tidy temporal tibble”) is a special data frame for time series. It extends the tibble data structure by formally declaring one column as the index (the time variable) and optionally, one or more columns as key(s) (unique identifiers for different series)

4.1.1 Key terminologies

Index: The variable (usually a date/time) that defines the time order of observations.

Key: A variable (or combination of variables) that uniquely identifies different time series’ within a single table e.g. country, stock_symbol etc.

Interval: The regular frequency of the measurements e.g. daily, monthly, quarterly etc.

To illustrate what we just discussed below is a standard dataframe with a date column. This data is not time aware yet (not a tsibble).

# A tibble: 4 × 3
  date       product sales
  <date>     <chr>   <dbl>
1 2023-01-01 A         120
2 2023-01-02 A         145
3 2023-01-01 B          88
4 2023-01-02 B         102

a tsibble on the other hand is self-aware. It knows its index and its key, unlocking powerful analysis tools. Compare the table below. The “[1D]” beside the tsibble’s dimension represents the interval (daily) and the “[2]” represents the number of unique keys/series.

# A tsibble: 4 x 3 [1D]
# Key:       product [2]
  date       product sales
  <date>     <chr>   <dbl>
1 2023-01-01 A         120
2 2023-01-02 A         145
3 2023-01-01 B          88
4 2023-01-02 B         102

4.2 Creating and Converting Data into a tsibble

creating a tsibble follows the same procedure as creating a normal data frame using the tsibble() function or simply from an existing data frame or tibble using the as_tsibble() function.

As explained earlier, these two functions require you to specify at least one of these two things (index or key) or both depending on your data.

We see how to create a tsibble from using tsibble() and as_tsibble below.

# single time series
tsibble(
  year = 2010:2050,
  value = rnorm(41),
  index = year
)

The above code creates a simple time series data using the tsibble() function, exactly like creating a dataframe. Indicating the index variable distinguishes a tsibble from a tibble.

Now we will create a tsibble from the as_tsibble() function using an existing tibble

# creating a tsibble from monthly data sales
# sales come from two different shops
sales_data <- tibble(
  Date = ymd(c('2025-01-01','2025-02-01','2025-03-01','2026-04-01','2025-05-01','2025-01-01','2025-02-01','2025-03-01','2025-04-01','2025-05-01')),
  Store = rep(c('Phone Shop', 'Beauty Shop'), each = 5),
  Sales = c(225, 150, 130, 90, 220, 190, 145, 180, 110, 180)
)
print(sales_data)
# We use as_tsibble to convert the tibble to a tsibble
# the date column becomes the index and the store column the key
sales_data_ts <- sales_data |> 
  mutate(Date = yearmonth(Date)) |>
  as_tsibble(
    index = Date,
    key = Store
  )
print(sales_data_ts)

We now have our sales data in a time aware tsibble format. The yearmonth() is an index function used to represent our Date in a year-month format.

4.2.1 Practice tsibble conversion

Here we will learn how to convert real world data into a tsibble. The data we will use is a time series data for Ghana containing certain world bank indicators. The data can be found here gh_data.csv. The data comes in a wide format (where years are columns). To make it tidy, where each row represents one indicator so that we have year is one column and value is another column, we first need to reshape it into a long format.

We will use pivot_longer from the tidyr package to reshape the data.

# load data
gh_data_raw <- read_csv('data/gh_data.csv', show_col_types = FALSE) 
  
# convert wide data to long
gh_data_long <- gh_data_raw |> 
  pivot_longer(
    cols = starts_with(c('19','20')), # select all date columns
    names_to = "Year",                # New column for years
    values_to = 'Value'               # New column for values
  )

# View the first few rows
head(gh_data_long)
# A tibble: 6 × 5
  `Country Name` `Indicator Name` `Indicator Code` Year    Value
  <chr>          <chr>            <chr>            <chr>   <dbl>
1 Ghana          Population_total SP.POP.TOTL      1960  6961215
2 Ghana          Population_total SP.POP.TOTL      1961  7162667
3 Ghana          Population_total SP.POP.TOTL      1962  7337375
4 Ghana          Population_total SP.POP.TOTL      1963  7514714
5 Ghana          Population_total SP.POP.TOTL      1964  7695739
# ℹ 1 more row

Success! Now each row is an observation of an indicator in a given year.

The next step is to convert the Year to a Proper Date. Right now Year is a character. We need it as a date so tsibble can understand time ordering. we will use the lubridate package functions and convert it to a Date object (assuming January 1st of each year).

# convert year to date
gh_data_long <-  gh_data_long |> 
  mutate(
    Year = as.integer(Year),                       # convert to integer number
    Date = lubridate::ymd(paste0(Year, "-01-01")), # Create date:Jan 1 of each year
    .after = Year                                  # add new Date column after year column
  )

# check structure
str(gh_data_long$Date)
 Date[1:1430], format: "1960-01-01" "1961-01-01" "1962-01-01" "1963-01-01" "1964-01-01" ...

Now Date is a proper Date object essential for time series analysis.

Finally we can convert our data to a time series dataframe (tsibble) that is time aware - index (time) and keys (unique series identifiers). In our case the Date column is our index and the Indicator_name becomes our key since we have multiple indicator names like population, life expectancy etc.

# convert data to tsibble
gh_ts <- gh_data_long |> 
  as_tsibble(
    index = Date,              # Time index
    key = `Indicator Name`     # Key column: each indicator is a separate series
  )

# View the tsibble
gh_ts
# A tsibble: 1,430 x 6 [1D]
# Key:       Indicator Name [22]
  `Country Name` `Indicator Name`       `Indicator Code`   Year Date       Value
  <chr>          <chr>                  <chr>             <int> <date>     <dbl>
1 Ghana          Annual GDP growth rate NY.GDP.MKTP.KD.ZG  1960 1960-01-01 NA   
2 Ghana          Annual GDP growth rate NY.GDP.MKTP.KD.ZG  1961 1961-01-01  3.43
3 Ghana          Annual GDP growth rate NY.GDP.MKTP.KD.ZG  1962 1962-01-01  4.11
4 Ghana          Annual GDP growth rate NY.GDP.MKTP.KD.ZG  1963 1963-01-01  4.41
5 Ghana          Annual GDP growth rate NY.GDP.MKTP.KD.ZG  1964 1964-01-01  2.21
# ℹ 1,425 more rows

We have our time series tsibble ready! But there are some nuances in the data, the first obvious ones are the column names - Country Name, Indicator Name and Indicator Code- we see that they are surrounded in back ticks (`) . This tells us that they do not follow the correct naming convention for variables in R (no spaces between words). The next one is the interval [1D]. The tsibble package thinks our data’s temporal interval is daily, because we have a full date with months and days. We will see how to fix this in the next chapter.