Moved 2026-10-04 (import done; do not archive until source PR #3 is closed/merged). Canonical copy lives in the kk-kb monorepo at
content/projects/01-vancouver-ai-community/special-features/hackathons/rival-2024-2025/round-4-music/vanai-hackathon-004/(source commit3cd45ed7bf74). Large media is in private Drive β see PROVENANCE.md there. This repo stays open for PR #3.
A dataset with 1,000+ survey responses about how people discover music, what formats they use, and how they feel about AI-generated music. Perfect for creating data visualizations, apps, or insights about music and technology.
Option 1: Download ZIP (Easiest)
- Click the green "Code" button at the top of this page
- Click "Download ZIP"
- Unzip the file on your computer
- Open
data/raw/music_survey_data.csvin Excel, Google Sheets, or any data tool
Option 2: Clone with Git
git clone https://github.com/WalksWithASwagger/vanai-hackathon-004.git
cd vanai-hackathon-004Option 3: Individual Files
- Right-click any file β "Save As" to download individual files
- Main dataset:
data/raw/music_survey_data.csv
- 1,000+ complete responses from music listeners across Canada
- 19 core questions covering music discovery, format evolution, and AI attitudes
- Rich demographics: Age, location, education, income, household composition
- 3,000+ text responses with sentiment analysis scores
- Geographic distribution: Ontario (40%+), BC (20%+), Alberta, Quebec, and more
- Music Discovery: How people find new music (radio, streaming, social media, etc.)
- Format Evolution: Personal experiences with vinyl, cassettes, CDs, downloads, streaming
- Listening Habits: When and how people consume music in daily life
- AI Music Attitudes: Responses to AI-generated music and voice cloning
- Social Sharing: How people share music and their guilty pleasures
- Personal Connection: Life theme songs and meaningful lyrics
- Open
data/raw/music_survey_data.csvin Excel or Google Sheets - Check out
data/raw/survey_questions.txtto see what was asked - Run
python scripts/explore_data.pyfor a quick data overview
Python Users
import pandas as pd
# Load the data
df = pd.read_csv('data/raw/music_survey_data.csv')
print(f"Total responses: {len(df)}")
# Check out some life theme songs
print(df['Q18_Life_theme_song'].dropna().head())R Users
library(tidyverse)
music_survey <- read_csv("data/raw/music_survey_data.csv")
glimpse(music_survey)Excel/Google Sheets Users Just open the CSV file - no coding required!
Q1_Relationship_with_music: Music engagement levelQ2_Discovering_music: Primary music discovery methodQ4_Music_format_changes: Experience with format transitionsQ7_New_music_discover_*: Multiple discovery methodsQ8_Music_listen_time_GRID_*: Listening habits by activityQ10_Songs_by_AI: Attitudes toward AI-generated musicQ11_Use_of_dead_artists_voice_feelings: AI voice cloning opinionsQ12_Music_bingo_*: Music-related behaviorsQ13_Share_the_music_you_love_*: Music sharing methodsQ15_Music_guilty_pleasure: Musical guilty pleasuresQ18_Life_theme_song: Personal theme songsQ19_Lyric_that_stuck_with_you: Meaningful lyrics
AgeGroup_Broad: 18-34, 35-54, 55+Province: Geographic distributionEducation: Education levelHH_Income_Fine_23: Household incomeGender: Gender distribution
- Sentiment scores included for open-ended responses
- Polarity and subjectivity measures available
- Percentage scores for emotional content
vanai-hackathon-004/
βββ π README.md # You are here!
βββ π data/
β βββ raw/
β β βββ music_survey_data.csv # π― THE MAIN DATASET (start here!)
β β βββ survey_questions.txt # What questions were asked
β βββ processed/ # Your cleaned data goes here
β βββ analysis/ # Your analysis results go here
βββ π§ scripts/
β βββ explore_data.py # Quick data peek script
β βββ data_preprocessing.py # Data cleaning helpers
βββ π» src/ # Your code goes here
βββ π― submissions/ # Hackathon projects go here
βββ π docs/ # Extra documentation
The Important Files:
data/raw/music_survey_data.csv- This is your main dataset!data/raw/survey_questions.txt- Explains what each question meansscripts/explore_data.py- Run this for a quick data overview
python scripts/explore_data.pyThis shows you basic info about the dataset - how many responses, what columns exist, sample data, etc.
python scripts/data_preprocessing.pyThis helps clean up responses, handle missing data, and prepare data for analysis.
Using this dataset, create an experience that's interactive, visual, or narrative-driven. The goal is to transform raw survey data into compelling insights about music, technology, and human behavior.
- Format Freedom: Visual display, interactive demo, live performance, or unconventional approach
- Bite-Sized Impact: Engaging experience within 2-5 minutes
- Technical Details: Include all software, hardware, and instructions needed
- Data Grounding: Story must be grounded in the actual data
- AI Integration: Incorporate AI tools to enhance your narrative
- Creativity & Innovation (25%): Fresh, original approach
- Clarity & Effectiveness (25%): Clear communication of insights
- Engagement & Impact (20%): Compelling and memorable experience
- Technical Execution (20%): Well-implemented solution
- Community Value (10%): Open-source tools, shared methodologies
- Dataset Released: September 15, 2025
- Submission Deadline: October 15, 2025
- Award Date: October 29, 2025
Your final folder of deliverables should include:
- PDF Document containing:
- Team member information (names, contact information)
- Project title and description (max 300 words)
- Brief explanation of your technical approach and tools used (max 300 words)
- Link to working prototype/demo, or the file itself
- 3-minute video walkthrough (strongly encouraged, especially for interactive projects)
Suggested tools, not requirements:
- Cline, RooCode, Cursor, Claude Code - AI copilots to jam with
- v0 by Vercel, Lovable - interfaces at speed
- Udio, Suno, ElevenLabs - sound experiments
- Pika, Runway, Google Veo - video riffs
- p5.js, Processing, Observable - data paintings
- Seedream 4 (ByteDance), Midjourney, Stable Diffusion - visual generation
- Excel/Google Sheets: Just open the CSV file and start exploring
- Python: Install pandas (
pip install pandas) and use the code examples above - R: Install tidyverse (
install.packages("tidyverse")) and use the R code above
- Download: Click green "Code" button β "Download ZIP"
- Questions: Check the Issues tab or create a new issue
- Sharing: Fork this repository to make your own copy
- Look at
data/raw/survey_questions.txtto understand the questions - Run
python scripts/explore_data.pyto see what's in the data - Check the Issues tab for common questions
This dataset is provided for hackathon use. Please respect participant privacy and use data responsibly.
π― Ready to start? Download the data and explore 1,000+ music stories from across Canada!