Chatbots are currently one of the most effective solutions for answering Frequently Asked Questions (FAQs). Considering the critical importance of speed and accurate answering, this project explores the possibility of designing a highly efficient chatbot for processing new, unseen questions. The primary objective of this research is to discover how the integration of additional contextual information and documentation affects the overall performance of a machine learning chatbot.
This project examines a fundamental trade-off in natural language processing: Which approach yields better results—a highly complex generative model, or a simpler model trained with richer contextual knowledge?.
To test this hypothesis, we conducted a comparative analysis between two distinct transformer architectures:
- GPT-2: Representing the more complex, generative model architecture.
- DistilBERT: Representing a simpler, lighter model trained not only on isolated questions and answers but also explicitly provided with the surrounding context of those questions.
We utilized SQuAD 2.0 (The Stanford Question Answering Dataset), a robust reading comprehension dataset consisting of questions posed by crowd-workers on a diverse set of Wikipedia articles.
- The dataset contains over 100,000 question-answer pairs spanning more than 500 articles.
- Crucially, it includes over 50,000 unanswerable questions to test the model's ability to recognize when sufficient information is missing.
- The answer to every valid question is an exact segment of text, or "span," extracted directly from the corresponding reading passage.
- The data was split into training and validation sets utilizing an 8-to-2 ratio.
DistilBERT is a small, fast, and lightweight model trained by distilling the BERT (Bidirectional Encoder Representations from Transformers) base architecture.
- Architecture Benefits: It features 40% fewer parameters than BERT and runs 60% faster, while still preserving over 95% of the original model's performance.
- Tokenization: We utilized the
DistilBertTokenizerFastfor tokenizing the text, truncating inputs to a maximum length of 512 tokens. - Fine-tuning: The model was fine-tuned specifically on the SQuAD 2.0 dataset, relying on the combined inputs of the
questions,context, andanswerscolumns. - Optimization: The network was optimized using the AdamW algorithm with a learning rate of 3e-4.
GPT-2 (Generative Pre-trained Transformer 2) is pre-trained on BookCorpus (a dataset of over 7,000 unpublished fiction books) and 8 million web pages.
- Architecture: We utilized the pre-trained GPT-2 small variant, which contains 124M parameters.
- Fine-tuning: Fine-tuned on SQuAD 2.0 using only the
questionsandanswerscolumns (omitting the rich context). - Optimization: Trained for 3000 steps with a learning rate of 0.001. Accuracy was calculated by evaluating similar word generation.
The empirical results of our fine-tuning process demonstrated a massive disparity in performance based on the presence of contextual data.
| Metric | DistilBERT (Context Provided) | GPT-2 (No Context Provided) |
|---|---|---|
| Accuracy / Precision | 85.30% | 0.11475 |
| Average Loss / F1 | 0.4028 | 0.11570 (F1 Score) |
Note: DistilBERT achieved its peak accuracy of 85.30% by the 6th epoch of training.
The experimental data conclusively demonstrates that a simpler model provided with richer, contextually relevant information heavily outperforms a significantly more advanced generative model operating with limited textual context.
- Limitations: The scope of this research was primarily constrained by time and computational resources. Due to the inherent limitations of the Google Colab GPU environment, we experienced occasional system crashes and were unable to scale up the models to utilize higher parameter counts or extended training epochs.
- Future Direction: Future iterations of this project will aim to increase overall performance by utilizing distributed, multi-GPU computing environments, exploring even more sophisticated datasets, and fine-tuning the simpler contextual models using advanced techniques such as Reinforcement Learning (RL) or Graph Neural Networks (GNN).