A Python-based web scraper to automatically download PDF files from Moodle courses at Frankfurt University of Applied Sciences.
✨ Created entirely with GitHub Copilot
- 🔐 Microsoft SSO Authentication: Handles Microsoft Single Sign-On with manual browser login
- 📚 Multi-Course Support: Download from multiple courses in one run
- 🤖 Browser Automation: Uses Selenium WebDriver to navigate Moodle
- 📥 Smart PDF Detection: Automatically finds and downloads PDF files
- 🗂️ Organized Storage: Creates separate folders for each course with proper names
- ⚡ Session Reuse: Extracts browser cookies for efficient downloads
- 🔄 Automatic Redirects: Follows Moodle's redirect mechanism to actual PDF files
- Python 3.9 or higher
- Chrome browser installed
- Active Frankfurt University Moodle account
-
Clone the repository
git clone https://github.com/yourusername/MoodleScraper.git cd MoodleScraper -
Create virtual environment
python3 -m venv .venv source .venv/bin/activate # On macOS/Linux # or .venv\Scripts\activate # On Windows
-
Install dependencies
pip install -r requirements.txt
-
Run the scraper
python moodlescraper.py
-
Manual Login
- Chrome browser will open automatically
- Log in with your Frankfurt University credentials
- Complete Microsoft 2FA if prompted
- Wait until Moodle dashboard loads
- Return to terminal and press Enter
-
Automatic Download
- Script will process all configured courses
- PDFs are downloaded to
~/Downloads/folder - Each course gets its own subfolder
Edit the COURSES dictionary in moodlescraper.py to customize which courses to download:
COURSES = {
7557: "Course Name 1",
218: "Course Name 2",
# Add more courses here
}Course IDs can be found in the Moodle URL: https://campuas.frankfurt-university.de/course/view.php?id=XXXXX
~/Downloads/
├── Course Name 1/
│ ├── Lecture_1.pdf
│ ├── Lecture_2.pdf
│ └── ...
├── Course Name 2/
│ ├── Exercise_1.pdf
│ └── ...
- Browser Automation: Selenium WebDriver launches Chrome and navigates to Moodle
- Authentication: User manually logs in with Microsoft SSO (bypasses automation detection)
- Section Expansion: JavaScript expands all collapsed course sections
- Resource Discovery: Finds all
mod/resource/view.phplinks on course pages - Cookie Transfer: Extracts session cookies from browser to requests library
- Redirect Following: Requests each resource URL, follows redirects to actual PDF files
- Smart Download: Checks content-type headers, only downloads PDFs
- Selenium WebDriver: Browser automation
- webdriver-manager: Automatic ChromeDriver management
- requests: HTTP downloads with session support
- urllib3: SSL handling
- SSL Warnings: LibreSSL 2.8.3 may show warnings (can be safely ignored)
- HTML Resources: Some Moodle resources are HTML pages, not PDFs (skipped automatically)
- Non-PDF Files: Word documents and other file types are skipped
Contributions are welcome! Feel free to:
- Report bugs
- Suggest new features
- Submit pull requests
MIT License - see LICENSE file for details
Created entirely with GitHub Copilot - AI pair programmer by GitHub
Developed for students at Frankfurt University of Applied Sciences to streamline course material downloads.
This tool is for educational purposes. Please respect your institution's terms of service and only download materials you have legitimate access to. The authors are not responsible for misuse of this tool.
Happy Learning! 📚✨