Day 7: Rise of the Pre-trained Giants - BERT
Day 7: Rise of the Pre-trained Giants - BERT
Objective
Understand the concept of pre-training and fine-tuning, and study the architecture of BERT, a model that revolutionized how we solve a wide range of NLP tasks.
Core Concepts
-
The Paradigm Shift: Pre-training and Fine-Tuning
- Pre-training: Train a very deep model (like a Transformer encoder) on a massive, unlabeled text corpus (e.g., all of Wikipedia). The model is not trained on a specific task, but on a general "language understanding" objective. This step is computationally very expensive.
- Fine-tuning: Take the pre-trained model, add a small, task-specific classification layer on top, and train it on a much smaller, labeled dataset for your actual task (e.g., sentiment analysis, question answering). The weights of the pre-trained model are only slightly adjusted ("fine-tuned").
-
BERT (Bidirectional Encoder Representations from Transformers):
- BERT's key innovation was to apply the bidirectional training of Transformers to language modeling.
- Unlike previous models that read text either left-to-right or right-to-left, BERT reads the entire sequence of words at once. It uses a Transformer encoder stack.
-
BERT's Pre-training Objectives:
- Masked Language Model (MLM):
- This is how BERT achieves bidirectionality. Instead of predicting the next word, it takes an input sentence and masks out about 15% of the words.
[CLS] The cat sat on the [MASK].- The model's objective is to predict the original identity of the masked words, using the context from both left and right.
- Next Sentence Prediction (NSP):
- The model receives two sentences, A and B, and must predict whether sentence B is the actual sentence that follows sentence A in the original text, or just a random sentence.
- This was intended to help the model understand sentence relationships, though later research found it to be less impactful than MLM.
- Masked Language Model (MLM):
🧠 Math & Stats Focus: The Softmax Function in Detail
BERT's output for a masked token is a probability distribution over the entire vocabulary. This is produced by a final Dense layer followed by a softmax activation.
- Softmax Function: Converts a vector of
Kreal numbers (logits) into a probability distribution ofKpossible outcomes.softmax(z)_i = e^(z_i) / Σ(e^(z_j))forj=1 to K - Properties:
- Outputs Probabilities: Each output value is between 0 and 1.
- Sums to 1: The sum of all output values is exactly 1.
- Highlights the Maximum: The exponentiation exaggerates the differences between the input logits. The largest logit value will get a much higher probability than the others, making the softmax output a good representation of a "choice".
- Cross-Entropy Loss: The softmax output is then compared to the true label (the one-hot encoded vector of the actual masked word) using cross-entropy loss, which we discussed on Day 2.
📜 Key Research Paper
- Paper: "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding" (Devlin et al., 2018)
- Link: https://arxiv.org/abs/1810.04805
- Contribution: BERT was a landmark achievement. It obtained state-of-the-art results on a wide range of NLP benchmarks (like GLUE and SQuAD) simply by fine-tuning the same pre-trained model. It demonstrated the immense power of deep, bidirectional pre-training and established this new paradigm for NLP.
💻 Project: Use a Pre-trained BERT for Sentiment Analysis
The power of BERT is that you don't need to pre-train it yourself. You can use it out-of-the-box for fine-tuning. The transformers library by Hugging Face makes this incredibly easy.
- Install Hugging Face:
pip install transformers datasets. - Load a Dataset: Use the
datasetslibrary to load a sentiment analysis dataset, likeimdb. - Load a Pre-trained BERT and Tokenizer:
from transformers import AutoTokenizer, AutoModelForSequenceClassificationtokenizer = AutoTokenizer.from_pretrained('bert-base-uncased')model = AutoModelForSequenceClassification.from_pretrained('bert-base-uncased', num_labels=2)
- Tokenize the Data: Use the tokenizer to convert your text sentences into the input format BERT expects (input IDs, attention mask).
- Fine-tune: Use the
TrainerAPI from Hugging Face or a standard PyTorch/TensorFlow training loop to fine-tune the model on the IMDB dataset. - Evaluate: See how well the fine-tuned model performs on the test set. You should be able to achieve very high accuracy (>90%) with minimal code.
Hugging Face provides excellent tutorials for this exact task. Following their Fine-tuning a pre-trained model guide is the best way to complete this project.
✅ Progress Tracker
- I can describe the pre-training/fine-tuning paradigm.
- I can explain what makes BERT "bidirectional" (the MLM objective).
- I understand the role and properties of the softmax function.
- I have successfully fine-tuned a pre-trained BERT model on a sentiment analysis task.