Automatically converts any YouTube lecture into structured academic notes in Markdown and PDF — just provide a URL.
project/
│
├── main.py # Entry point — runs the full pipeline
├── config.py # Path & title configuration, initialized once
├── transcript.py # YouTube transcript extraction (captions or Whisper)
├── ai_agent_summarizer.py # LLM-based summarization via Groq
├── markdown_to_pdf.py # Converts markdown output to PDF via pandoc
├── header.tex # LaTeX styling for PDF export
├── .env # API keys (not committed)
│
└── Outputs/
└── <Video Title>/
├── Transcript.txt
├── AI_Summary.md
└── AI_Summary.pdf
The pipeline runs in 3 stages when you provide a YouTube URL:
YouTube URL
│
▼
[1] transcript.py → Extracts transcript (captions or Whisper fallback)
│
▼
[2] ai_agent_summarizer.py → Sends transcript to Groq LLaMA-3.3-70b
│ Generates structured academic markdown notes
▼
[3] markdown_to_pdf.py → Converts markdown to PDF via pandoc + xelatex
- Set your YouTube URL in
main.py:
YOUTUBE_URL = "https://youtu.be/your_video_id"- Run the pipeline:
python main.pyOutput files are saved to Outputs/<Video Title>/.
pip install -r requirements.txtCreate a .env file in the project root:
GROQ_API_KEY=your_groq_api_key_here
Get a free Groq API key at console.groq.com.
The generated notes follow an academic structure:
- Numbered headings and subheadings
- Bullet-point explanations for each topic
- Expanded technical details for briefly mentioned concepts
- Properly formatted code snippets
- Comparison tables where applicable
- Full markdown formatting, PDF-ready
transcript.py tries two methods automatically:
| Method | When Used |
|---|---|
| YouTubeTranscriptAPI | Video has captions available |
OpenAI Whisper (base model) |
No captions found — downloads audio and transcribes locally |
PDF output is styled via header.tex using:
- 1-inch margins
- 1.2x line spacing
- Monospaced code blocks with a light grey background (
DejaVu Sans Mono)
To customize, edit header.tex directly.
- Re-running the pipeline on the same URL will overwrite all previous output files automatically.
config.pymust be initialized viaconfig.init(url)before any other module is imported — this is handled bymain.py.markdown_to_pdf.pyusesheader.texresolved relative to the script's location, so it works regardless of where you runmain.pyfrom.