A Study of the HarvardX-MITx Person-Course Dataset with Discrete Choice Models
Since the beginning of 2020, COVID-19 has developed into a pandemic, influencing millions of people around the world. Surely the pandemic has interrupted people's life, especially for students who rely heavily on face-to-face learning. However, as the pandemic has accelerated the digital transformation, online education has been made more easily assessed, offering another opportunity. Based on the data from China Internet Network Information Center, by June 2019, there are 854 million netizens in China. Among them, 232 million use online learning. During the pandemic, 265 million students moved to online, boosting both user base and frequency of use. By March 2020, the user base has reached 423 million rapidly and the use frequency increases from 27.2% to 46.8%. Meanwhile, since many colleges like Tsinghua offer free online courses to the public, online learning became even more popular.
One popular online learning form is MOOC. MOOC is the abbreviation of massive open online course, which is an online course aimed at unlimited participation and open access via the web (Kaplan et al., 2016). It was first introduced in 2008 and emerged as a popular mode of learning in 2012 (Siemens, 2013). MIT and Harvard, as university pioneers, joined in the initiative. In the year from the fall of 2012 to the summer of 2013, the first 17 HarvardX and MITx courses launched on the edX platform (Ho et al., 2014). During the past decade, millions of students from around the globe have enrolled; thousands of courses have been offered; and hundreds of universities have joined hands to revolutionize education. Therefore, it is vital to understand people's choice behaviors concerning online education so that we can better predict demands for online learning and offer attractive courses.
Prior researches have been focusing on MOOCs as a phenomenon and perform quantitative analysis. One possible reason could be a lack of data since MOOCs platforms like edX may rely on data for profit. Privacy concerns could be another reason. One qualitative research was conducted by Christensen et al. in 2014. They used the survey data based on students enrolled in the University of Pennsylvania’s MOOCs to analyze who takes MOOCs and why they take MOOCs. Based on stated preferences data, they found that students' main reasons for taking a MOOC are advancing in their current job and satisfying curiosity (Christensen et al., 2014). In our research, we use the HarvardX-MITx Person-Course Dataset, which involves revealed preferences data collected on the edX platform to further analyze people's choice behavior in terms of MOOCs. Specifically, we want to understand how people choose from multiple courses and how they choose to get certificates. As far as we know, we are the first to use discrete choice models to analyze the problem.