5  Dealing with Time Gaps and Irregularities

Real-world data is messy. Sensors fails, public holidays happen, data does not get entered. This leads to gaps in our time series – missing entries in the index where we expect a measurement. Traditional data frames ignore this, but tsibble helps you find and handle them gracefully.

5.1 Why Time Gaps Matter

Many time series models and visualisations assume regular spaced data. Gaps can lead to errors, misleading plots and inaccurate forecasts.

5.2 How tsibble Helps

We can check the time gaps in our tsibble using the scan_gaps() function from the tsibble package. This is a very handy function that compares our actual data against a complete regular timeline and tells us exactly what is missing. Other equally useful functions for handling time gaps include count_gaps() and has_gaps().

We will scan our gh_ts tsibble to check if there are gaps. But before that, we will fix the column names with janitor::clean_names() to avoid complications later on in our analysis due to the unconventional column names.

gh_ts <- gh_ts |> 
  janitor::clean_names() 

gh_ts |> scan_gaps()
# A tsibble: 512,864 x 2 [1D]
# Key:       indicator_name [22]
  indicator_name         date      
  <chr>                  <date>    
1 Annual GDP growth rate 1960-01-02
2 Annual GDP growth rate 1960-01-03
3 Annual GDP growth rate 1960-01-04
4 Annual GDP growth rate 1960-01-05
5 Annual GDP growth rate 1960-01-06
# ℹ 512,859 more rows

the scan_gaps() function returns a long list of missing dates. However, this is not true for our data set. We are getting this many gaps because the inner workings of tsibble thinks that the days of all the months are missing, meanwhile our data is not a daily series but a yearly series. This whole misinterpretation comes from the incorrect interval [1D]. To fix this we can change the date to a year-month format or use the year column as the index and then rescan for gaps. We will use both options to demonstrate the interval change.

# fix interval  by assigning index to year
gh_ts |> 
  as_tsibble(
    index = year
  ) |> head(2)
# A tsibble: 2 x 6 [1Y]
# Key:       indicator_name [1]
  country_name indicator_name         indicator_code     year date       value
  <chr>        <chr>                  <chr>             <dbl> <date>     <dbl>
1 Ghana        Annual GDP growth rate NY.GDP.MKTP.KD.ZG  1960 1960-01-01 NA   
2 Ghana        Annual GDP growth rate NY.GDP.MKTP.KD.ZG  1961 1961-01-01  3.43

Here we see that the interval is now correctly represented as [1Y] which is exactly what we expect. We could have equally modified the date column with the yearmonth() function and kept it as the index without having to change the index to the year column.

gh_ts <- gh_ts |> 
  mutate(date = yearmonth(date))

gh_ts |> head(3)
# A tsibble: 3 x 6 [12M]
# Key:       indicator_name [1]
  country_name indicator_name         indicator_code     year     date value
  <chr>        <chr>                  <chr>             <dbl>    <mth> <dbl>
1 Ghana        Annual GDP growth rate NY.GDP.MKTP.KD.ZG  1960 1960 Jan NA   
2 Ghana        Annual GDP growth rate NY.GDP.MKTP.KD.ZG  1961 1961 Jan  3.43
3 Ghana        Annual GDP growth rate NY.GDP.MKTP.KD.ZG  1962 1962 Jan  4.11

Our interval for this new modification is [12M] - meaning 12 months, essentially indicating a year. We can go ahead and check for the gaps now.

gh_ts |> scan_gaps()
# A tsibble: 0 x 2 [?]
# Key:       indicator_name [0]
# ℹ 2 variables: indicator_name <chr>, date <mth>

This returns an empty tsibble, telling us that there are no time gaps. we can confirm again with the has_gaps() function.

gh_ts |> has_gaps() |> 
  pull(.gaps) |> sum()
[1] 0

Confirms Zero Gaps in our gh_ts data!

5.3 Handling Time Gaps

In order for us to understand the concept of time gaps and how to handle them properly we will simulate a data with implicit time gaps and then see how to deal with them using functions from the *_gap() functions.

We will use the same product sales concept to simulate this series. Here we compare sales for 2 products Smart Phone and Laptop over a 12 month period.

# create a sequence of sales dates
start_date <- ym('2025-01')  
end_date <- ym('2025-12')
sales_period <- seq.Date(
  from = start_date,
  to = end_date,
  by = 'month'
 ) |> 
  yearmonth()     # change format to year-month

# simulate data with missing dates for different products
sales_data_gaps <- tsibble(
  Product = c(rep('Smart Phone', 10), rep('Laptop', 8)),
  Sales = round(c(rnorm(10,300,65), runif(8, 620, 1000))),
  Date = c(
    sales_period[c(1:5,8:12)],      # Smart phone is missing Jun & Jul
    sales_period[c(2,3,4,5,6,9:11)] # Laptop is missing Jan, Jul, Aug & Dec
  ),
  index = Date,
  key = Product
)

print(sales_data_gaps)
# A tsibble: 18 x 3 [1M]
# Key:       Product [2]
  Product Sales     Date
  <chr>   <dbl>    <mth>
1 Laptop    771 2025 Feb
2 Laptop    923 2025 Mar
3 Laptop    699 2025 Apr
4 Laptop    670 2025 May
5 Laptop    968 2025 Jun
# ℹ 13 more rows

Our simulated tsibble (sales_data_gaps) is now ready. A visual inspection of the Date column will reveal that there are missing dates (implicit). we can confirm this with the functions we have learnt earlier and then decide what to do later.

Note

More information on the seq.Date function in (Chapter 7)

scan_gaps(sales_data_gaps)
# A tsibble: 4 x 2 [1M]
# Key:       Product [2]
  Product         Date
  <chr>          <mth>
1 Laptop      2025 Jul
2 Laptop      2025 Aug
3 Smart Phone 2025 Jun
4 Smart Phone 2025 Jul

The scan gaps tells us the gaps in our data. Notice how it is only showing that 2 months are missing for Laptop even though we know there are rather 4 months missing. We will see how to handle this very soon.

Important

The tsibble assumes that our data only starts from Feb and Ends in Nov for the Laptop Product.

The tsibble package makes it really easy to handle gaps in our time series data with the fill_gaps() function. The fill_gaps() makes the implicit gaps explicit by inserting rows for missing time periods and assigning NA to the observation values (the Sales column in our case).

# fill in the missing time points 
sales_data_gaps |> 
  fill_gaps() |> pull(Date)
<yearmonth[22]>
 [1] "2025 Feb" "2025 Mar" "2025 Apr" "2025 May" "2025 Jun" "2025 Jul"
 [7] "2025 Aug" "2025 Sep" "2025 Oct" "2025 Nov" "2025 Jan" "2025 Feb"
[13] "2025 Mar" "2025 Apr" "2025 May" "2025 Jun" "2025 Jul" "2025 Aug"
[19] "2025 Sep" "2025 Oct" "2025 Nov" "2025 Dec"

The fill_gaps() function filled only the gaps which were identified by the scan_gaps() function evidenced by only 22 time points instead of 24.

So you might be wondering, how do we deal with the missing Jan and Dec. Well! the fill_gaps() function creates provision for such cases. The function can even help us set a start/ending time that allows us to expand the existing time span (we are not going to do that).

Back to our issue, we can set .full = TRUE inside fill_gaps() to fill the time gaps over the entire time span of our data. What it does is, it checks the time span for the Smart Phone’s series and notices that it starts from Jan and ends in Dec so it applies the same time span to Laptop.

# fill gaps over entire time span
sales_data_filled <- sales_data_gaps |> 
  fill_gaps(.full = TRUE)

Also, instead of assigning the values of the filled time gaps with NAs you can specify a value of our choosing which is very useful if you want to record zero (0) sales for the missing months instead of missing data.

# fill gaps in Sales column with 0
sales_data_gaps |> 
  fill_gaps(
    .full = TRUE,
    Sales = 0
  )

Now, all missing time points have a Sales value of 0 instead of NA. This can be crucial for consistent time series modelling or calculations. Sometimes too, the data might not have time gaps but rather missing values for a particular time period. We can deal with them by filtering out complete cases (which automatically introduces time gaps) or use models that can handle NAs within a time series data. The former choice would not help much. Since we will not go into imputing missing data my advice is to rely on the fill_gaps() and replacing NAs with zeros.

Tip

the imputeTS package provides several robust methods for estimating missing values in a time series data.

Caution

Always apply domain knowledge and check time series data characteristics before applying a particular imputation method from imputeTS

The next chapter is a bonus one. It briefly describes how to import data (time series data) into R from various sources with a simple code illustration. You can decide to skip it if you already know how to import the data from some common file formats like .csv or .xlsx (excel file) into R.