Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

🛍️ Amazon Product Review Sentiment Analysis

A deep learning NLP project that classifies Amazon product reviews as Positive or Negative using an LSTM-based neural network trained on 1 million reviews from the Amazon Review Polarity Dataset.


📋 Table of Contents


Overview

This project performs binary sentiment classification on Amazon product reviews. Reviews with ratings of 1–2 stars are labelled Negative and reviews with ratings of 4–5 stars are labelled Positive (3-star reviews are excluded as neutral). The model uses word embeddings followed by an LSTM layer to capture sequential context in text, trained end-to-end on 1,000,000 review samples.


📦 Dataset

Amazon Reviews Polarity Dataset — Version 3 (Updated 09/09/2015)

Dataset Link : https://www.kaggle.com/datasets/bhavikardeshna/amazon-customerreviews-polarity

Originally compiled from ~35 million Amazon reviews spanning 18 years (up to March 2013) by J. McAuley and J. Leskovec, and reformatted into a polarity benchmark by Xiang Zhang for the NIPS 2015 paper "Character-level Convolutional Networks for Text Classification".

Split Negative Samples Positive Samples Total
Train 1,800,000 1,800,000 3,600,000
Test 200,000 200,000 400,000

CSV Format — each file has 3 columns:

Column Description
class 1 = Negative, 2 = Positive
title Review headline
text Full review body text

This project trains on the text column (review body) of the first 1,000,000 training samples, with an 80/20 train/validation split.


🧠 Model Architecture

A sequential deep learning model built with Keras:

Input (variable-length padded token sequences)
        │
        ▼
┌──────────────────────────────────┐
│  Embedding Layer                 │  vocab_size × 100 dims
└──────────────────────────────────┘
        │
        ▼
┌──────────────────────────────────┐
│  Batch Normalization             │
└──────────────────────────────────┘
        │
        ▼
┌──────────────────────────────────┐
│  LSTM (32 units)                 │  return_sequences=False
└──────────────────────────────────┘
        │
        ▼
┌──────────────────────────────────┐
│  Batch Normalization             │
└──────────────────────────────────┘
        │
        ▼
┌──────────────────────────────────┐
│  Dense (1 unit, sigmoid)         │  Binary output
└──────────────────────────────────┘
        │
        ▼
    Output: P(Positive Review)
Component Detail
Embedding 100-dimensional word vectors, trained from scratch
LSTM 32 hidden units, processes full sequence
Batch Normalization Applied after Embedding and LSTM for training stability
Output Activation Sigmoid (binary classification)
Loss Function Binary Cross-Entropy
Optimizer Adam
Metrics Accuracy, AUC
Batch Size 4096
Epochs 10

🔄 Pipeline

Raw CSV Data
    │
    ▼
1. Load train.csv (3.6M rows) → sample 1,000,000 reviews
    │
    ▼
2. Extract features (X = review text) and labels (y = class)
   Labels shifted: 1,2 → 0,1 (binary)
    │
    ▼
3. Train/Validation Split (80% train / 20% val, random_state=42)
    │
    ▼
4. Tokenization
   - Keras Tokenizer with OOV token
   - Fit on training text only
   - Build vocabulary (word → integer index)
    │
    ▼
5. Sequence Encoding
   - texts_to_sequences (train, val, test)
   - pad_sequences (post-padding to max_len)
    │
    ▼
6. Model Training (10 epochs, batch_size=4096)
   - Validation on val set each epoch
    │
    ▼
7. Model Saved to Google Drive (.h5)
    │
    ▼
8. Evaluation on test.csv (400,000 reviews)
   - Test Accuracy & Test AUC reported

📊 Results

The model is evaluated on the held-out test set of 400,000 reviews:

Metric Value
Test Accuracy Printed at runtime
Test AUC Printed at runtime

To see exact numbers, run the notebook end-to-end. The model outputs P(Positive) per review; a threshold of 0.5 is used for the binary prediction.


🛠️ Tech Stack

Library Purpose
Python 3 Core language
TensorFlow / Keras Model building and training
NumPy Numerical operations
pandas Data loading and manipulation
scikit-learn Train/validation split
Google Colab Cloud GPU training environment
Google Drive Dataset storage and model saving

📁 Project Structure

Amazon-Review-Sentiment/
│
├── product_review_analysis.ipynb   # Main Jupyter notebook (full pipeline)
├── readme.txt                      # Original dataset description
│
├── (external — on Google Drive)
│   ├── train.csv                   # 3.6M training reviews
│   ├── test.csv                    # 400K test reviews
│   └── model.h5                    # Saved trained model

🚀 Getting Started

This project is designed to run on Google Colab with data stored on Google Drive.

Step 1 — Upload Dataset to Google Drive

Download the Amazon Reviews Polarity Dataset and place both train.csv and test.csv in your Google Drive root (My Drive/).

Dataset source: Amazon Reviews Polarity — Fast.ai or Kaggle.

Step 2 — Open the Notebook in Colab

Upload product_review_analysis.ipynb to Google Colab, or open it directly from GitHub:

File → Open Notebook → GitHub → notstanzinn/Product-Review-Analyzer

Step 3 — Mount Google Drive

The notebook mounts your Drive automatically:

from google.colab import drive
drive.mount('/content/drive')

Step 4 — Run All Cells

Use Runtime → Run All. Training takes approximately 10 epochs on a Colab GPU (T4 recommended for this dataset size).

⚡ Enable GPU: Runtime → Change Runtime Type → GPU

Step 5 — Evaluate

The final cells load the saved model and print:

Test Accuracy: XX.XX%
Test AUC: X.XX

By notstanzinn

About

product review analyzer using neural networks

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages