A deep learning NLP project that classifies Amazon product reviews as Positive or Negative using an LSTM-based neural network trained on 1 million reviews from the Amazon Review Polarity Dataset.
This project performs binary sentiment classification on Amazon product reviews. Reviews with ratings of 1–2 stars are labelled Negative and reviews with ratings of 4–5 stars are labelled Positive (3-star reviews are excluded as neutral). The model uses word embeddings followed by an LSTM layer to capture sequential context in text, trained end-to-end on 1,000,000 review samples.
Amazon Reviews Polarity Dataset — Version 3 (Updated 09/09/2015)
Dataset Link : https://www.kaggle.com/datasets/bhavikardeshna/amazon-customerreviews-polarity
Originally compiled from ~35 million Amazon reviews spanning 18 years (up to March 2013) by J. McAuley and J. Leskovec, and reformatted into a polarity benchmark by Xiang Zhang for the NIPS 2015 paper "Character-level Convolutional Networks for Text Classification".
| Split | Negative Samples | Positive Samples | Total |
|---|---|---|---|
| Train | 1,800,000 | 1,800,000 | 3,600,000 |
| Test | 200,000 | 200,000 | 400,000 |
CSV Format — each file has 3 columns:
| Column | Description |
|---|---|
class |
1 = Negative, 2 = Positive |
title |
Review headline |
text |
Full review body text |
This project trains on the
textcolumn (review body) of the first 1,000,000 training samples, with an 80/20 train/validation split.
A sequential deep learning model built with Keras:
Input (variable-length padded token sequences)
│
▼
┌──────────────────────────────────┐
│ Embedding Layer │ vocab_size × 100 dims
└──────────────────────────────────┘
│
▼
┌──────────────────────────────────┐
│ Batch Normalization │
└──────────────────────────────────┘
│
▼
┌──────────────────────────────────┐
│ LSTM (32 units) │ return_sequences=False
└──────────────────────────────────┘
│
▼
┌──────────────────────────────────┐
│ Batch Normalization │
└──────────────────────────────────┘
│
▼
┌──────────────────────────────────┐
│ Dense (1 unit, sigmoid) │ Binary output
└──────────────────────────────────┘
│
▼
Output: P(Positive Review)
| Component | Detail |
|---|---|
| Embedding | 100-dimensional word vectors, trained from scratch |
| LSTM | 32 hidden units, processes full sequence |
| Batch Normalization | Applied after Embedding and LSTM for training stability |
| Output Activation | Sigmoid (binary classification) |
| Loss Function | Binary Cross-Entropy |
| Optimizer | Adam |
| Metrics | Accuracy, AUC |
| Batch Size | 4096 |
| Epochs | 10 |
Raw CSV Data
│
▼
1. Load train.csv (3.6M rows) → sample 1,000,000 reviews
│
▼
2. Extract features (X = review text) and labels (y = class)
Labels shifted: 1,2 → 0,1 (binary)
│
▼
3. Train/Validation Split (80% train / 20% val, random_state=42)
│
▼
4. Tokenization
- Keras Tokenizer with OOV token
- Fit on training text only
- Build vocabulary (word → integer index)
│
▼
5. Sequence Encoding
- texts_to_sequences (train, val, test)
- pad_sequences (post-padding to max_len)
│
▼
6. Model Training (10 epochs, batch_size=4096)
- Validation on val set each epoch
│
▼
7. Model Saved to Google Drive (.h5)
│
▼
8. Evaluation on test.csv (400,000 reviews)
- Test Accuracy & Test AUC reported
The model is evaluated on the held-out test set of 400,000 reviews:
| Metric | Value |
|---|---|
| Test Accuracy | Printed at runtime |
| Test AUC | Printed at runtime |
To see exact numbers, run the notebook end-to-end. The model outputs P(Positive) per review; a threshold of 0.5 is used for the binary prediction.
| Library | Purpose |
|---|---|
| Python 3 | Core language |
| TensorFlow / Keras | Model building and training |
| NumPy | Numerical operations |
| pandas | Data loading and manipulation |
| scikit-learn | Train/validation split |
| Google Colab | Cloud GPU training environment |
| Google Drive | Dataset storage and model saving |
Amazon-Review-Sentiment/
│
├── product_review_analysis.ipynb # Main Jupyter notebook (full pipeline)
├── readme.txt # Original dataset description
│
├── (external — on Google Drive)
│ ├── train.csv # 3.6M training reviews
│ ├── test.csv # 400K test reviews
│ └── model.h5 # Saved trained model
This project is designed to run on Google Colab with data stored on Google Drive.
Download the Amazon Reviews Polarity Dataset and place both train.csv and test.csv in your Google Drive root (My Drive/).
Dataset source: Amazon Reviews Polarity — Fast.ai or Kaggle.
Upload product_review_analysis.ipynb to Google Colab, or open it directly from GitHub:
File → Open Notebook → GitHub → notstanzinn/Product-Review-Analyzer
The notebook mounts your Drive automatically:
from google.colab import drive
drive.mount('/content/drive')Use Runtime → Run All. Training takes approximately 10 epochs on a Colab GPU (T4 recommended for this dataset size).
⚡ Enable GPU:
Runtime → Change Runtime Type → GPU
The final cells load the saved model and print:
Test Accuracy: XX.XX%
Test AUC: X.XX
By notstanzinn