Skip to content

5: Design

tarrantrl edited this page Nov 28, 2017 · 3 revisions

Data Source

Instead of using an API, we scraped our data directly from iTunes. We used Scrapy to get information about songs, artists, and albums using the top 100 songs from iTunes. The code outputs a json file each for the artists, albums, and songs. For artists, we got the name of the artist, the iTunes link to the artist, the iTunes link to the artist's picture, the place of origin of the artist, the artist's genre, the date the artist was born (or the date the group formed for bands), the artist's latest release, and the biographical information about the artist. For albums, we scraped the name of the album, the iTunes link to the album, the iTunes link to the album cover image, the description of the album, the artist of the album, the genre of the album, the number of songs on the album, and the actual track listing for the album. For songs, we acquired the song name, the artist of the song, the iTunes daily ranking of the song, the iTunes link to the album the song was from (since songs are played from the album page on iTunes), the duration of the song, the genre of the song, and the date the song was released. Although we only scraped data once, this code could be used to scrape data from iTunes daily. For more details about our scraping process or to view the code, see https://github.com/SuperYuLu/itunes-Scrapy.

Using the Database

Running the app locally

To run the app locally, we used PostgreSQL 9.5.10, which is an open source object-relational database. We created a database in PostgreSQL, which can be done through the graphical pgAdmin interface or through the command line. Configuration with the local database can be seen in our file models.py, which sets the SQLALCHEMY_DATABASE_URI to use PostgreSQL if that variable is not already present in the environment. Creating and populating the tables was all done using SQLAlchemy as described in the tools section of this wiki.

Running the app on GCP

To deploy the app on GCP, we created a database through GCP's SQL storage option. We created an instance of a PostgreSQL database and then created a database within that instance. In order to configure the app with GCP, we set the environment variable SQLALCHEMY_DATABASE_URI based on our specific user, password, database instance, and database name. Using os.getenv in models.py, we check whether this variable is defined in the environment (which it only will be when deploying on GCP) and use this variable if so. As when running locally, creation and population of the tables was set up using SQLAlchemy in our Python files.

Search Functionality

For each of our pillars (artists, albums, and songs), we added the capability to search the model tables. The search can be done across all attributes or within specific attributes. The search returns matches for whole words or partial words; for example, searching "one" in the artist table returns the artist "Malone" since that name contains the string "one." We implemented the search functionality using an existing JavaScript package. In the file tableConfig.js configures the settings to use the packages jquery.tablesorter.widgets.js, jquery.tablesorter.pager.js, and jquery.tablesorter.widgets.js.bak (which are publicly available packages). In the HTML files for each model, the header references the configuration file and the pager package.

Clone this wiki locally