A GNN-based book recommendation system models users, books, and their interactions as a graph, leveraging graph neural networks to capture complex, high-order relationships for more accurate and personalized recommendations.
This project utilizes the Amazon Book Dataset. The raw data consists of the following files:
t10k-images-idx3-ubyte.gzt10k-labels-idx1-ubyte.gz
Before diving into model training, it is highly recommended to get familiar with the data distribution and content.
I have written a utility script for this purpose:
You can run datasetpreview.py to quickly preview the dataset details, including data volume, label mapping, and sample structures.
After becoming familiar with the dataset, you can start training your own recommendation engine.
This project adopts LightGCN, a lightweight yet powerful variant of Graph Neural Networks (GNN). By removing unnecessary feature transformations and non-linearities, LightGCN efficiently captures high-order collaborative filtering signals between users and books.
The trained recommendation model successfully generates personalized suggestions, providing the top 10 books with the highest confidence scores for every user.
To make the results intuitive, I have applied visual analytics to the recommendation outputs. You can check out the following visualization materials included in the repository:
The foundation of the system is how we define the graph
| Element | Description |
|---|---|
| Nodes ( |
Two main types: User nodes ( |
| Edges ( |
Interaction behaviors: rate, purchase, add_to_cart, like, review. |
| Node Features | User: age, location; Book: title/blurb embeddings (BERT), price, avg. rating. |
| Edge Weights | Implicit (0/1) or explicit (rating score, confidence level). |
Typically modeled as a Bipartite Graph, where edges only exist between
GNNs learn node representations by iteratively aggregating feature information from neighboring nodes.
For recommendation scenarios, simpler and more efficient models usually outperform complex ones:
- LightGCN (Highly Recommended)
- Removes feature transformation matrices and non-linear activation functions.
- Purely uses neighborhood aggregation to refine embeddings—lightweight and state-of-the-art for collaborative filtering.
- NGCF (Neural Graph Collaborative Filtering)
- Exploits high-order connectivities in user-item bipartite graphs explicitly.
- GraphSAGE / GAT
- Useful when incorporating rich side information (e.g., book abstracts) and needing attention mechanisms to weigh neighbor importance.
Stack
Combine embeddings from all layers (usually via weighted sum or concatenation):
Calculate the matching score between user
Use BPR (Bayesian Personalized Ranking) loss to ensure the predicted score of a positively interacted book is higher than an unobserved one: