Grants and Contributions:

Title:
Distributional effects and computational challenges in modern data analysis
Agreement Number:
RGPIN
Agreement Value:
$175,000.00
Agreement Date:
May 10, 2017 -
Organization:
Natural Sciences and Engineering Research Council of Canada
Location:
Ontario, CA
Reference Number:
GC-2017-Q1-03384
Agreement Type:
Grant
Report Type:
Grants and Contributions
Additional Information:

Grant or Award spanning more than one fiscal year. (2017-2018 to 2022-2023)

Recipient's Legal Name:
Volgushev, Stanislav (University of Toronto)
Program:
Discovery Grants Program - Individual
Program Purpose:

The research agenda described in this proposal has two main threads: procedures for inference in massive data sets and the use of copulas in time series analysis.

With modern data collection techniques, extremely large data sets become more and more prevalent. To extract useful information from such data sets, statistical procedures that can handle large amounts of data and remain computationally feasible need to be developed. This motivates the first research direction: fast bootstrap procedures for massive data sets. The goal of this part of the proposed research is to gain a deep and comprehensive understanding of the limitations associated with bootstrap procedures that are specifically designed for large-scale data sets. This will be achieved by analyzing a wide array of scenarios where the classical bootstrap or modifications thereof are known to be applicable. In cases where the available methods fail, I plan to develop alternative approaches that continue to be applicable. With data sets that are growing at an ever increasing rate, fast and accurate procedures for quantifying uncertainty of statistical procedures are in high demand. My long-term goal is to enable the scientific community to conduct fast and reliable inference for the new type of large and messy data sets that are becoming more and more common in the modern data world. The research described above will provide a first step towards a deep and comprehensive study of such procedures and enable researchers and industry professionals to make full use of the data they collect.

The second main topic of this proposal is the use of copulas in time series analysis. Since the financial crisis, it is well known that correct modeling of complex dependencies among economic variables or financial instruments is extremely important. Copulas provide a simple and elegant way to find and validate such models.
During the next five years, I aim to extend classical tools from time series analysis to allow visualization and analysis of distributional effects in dynamics of time series by using copulas. To this end I aim to provide the statistical community with a toolbox of methods that are grounded in solid theoretical understanding, well-documented and understandable to the non-technical time series and copula community, and have a fast and reliable implementation in R in order to facilitate the wide-spread use of such methods in a broad community of applied researchers and beyond.