An NLP-based classifier that flags fraudulent job postings and recruiter messages, with an interpretable breakdown of why a message looks like a scam — not just a black-box label.
Online job scams — fake "work from home" offers, upfront "registration fees," requests for bank details before any real interview — are common and costly. This project was born out of a real experience reporting a job scam, and turns that experience into a tool: paste in a job posting or recruiter message, and get back a scam probability plus the specific red flags that triggered it.
Most people spot scams only after losing money or sharing sensitive information. This tool gives a fast, explainable second opinion before you engage — combining a machine learning classifier with transparent, rule-based red-flag detection so the output is trustworthy and actionable, not just a score.
- Curated dataset — hand-collected scam and legitimate job posting examples covering common Indian and international scam patterns (upfront fees, urgency language, requests for bank/Aadhar details, informal-only contact channels) vs. normal hiring language.
- Feature engineering — beyond TF-IDF text vectors, extracts explicit red-flag signals: urgency score, payment-request score, sensitive-info-request score, informal-channel score, "no interview needed" score, exclamation/caps usage.
- Model — Logistic Regression over combined TF-IDF + scaled engineered features, with
class_weight="balanced"for the two classes. - Explainability layer — a separate rule-based
explain_flags()function surfaces which red flags fired in plain English, alongside the model's probability score. - Interactive app — a small Flask UI to paste in any message and get an instant check.
| Metric | Score |
|---|---|
| Accuracy (held-out split) | 1.00 |
| F1-score | 1.00 |
Python scikit-learn pandas Flask TF-IDF Logistic Regression
git clone https://github.com/naax005/job-scam-detector.git
cd job-scam-detector
pip install -r requirements.txt
# Train the model
python src/train.py
# Run the web app
python app.py
# then visit http://127.0.0.1:5000Or use it directly in Python:
from src.predict import predict
result = predict("Congratulations! Pay a $49 registration fee via Western Union to activate your work from home job today!!!")
print(result["label"], result["scam_probability"])
print(result["red_flags"])job-scam-detector/
├── data/
│ └── job_postings.csv # curated scam vs legit examples
├── src/
│ ├── feature_engineering.py # red-flag feature extraction + explanations
│ ├── train.py # trains TF-IDF + Logistic Regression model
│ └── predict.py # scores new text, returns probability + flags
├── models/ # saved model, vectorizer, scaler (generated)
├── templates/
│ └── index.html # Flask UI
├── app.py # Flask web app
├── requirements.txt
└── README.md
- Scale up the dataset using public scam-report datasets and real scraped job postings for better generalization
- Add a browser extension so people can check messages without leaving LinkedIn/WhatsApp
- Try a transformer-based model (DistilBERT) for better handling of subtle/ambiguous phrasing
- Add domain/URL reputation checks for links included in messages
- Track false positive rate specifically on legitimate postings that use urgent-sounding but normal hiring language
Alinas Ferdavus — LinkedIn
Built after personally encountering and reporting an online job scam via cybercrime.gov.in — this project is part of a broader interest in applying data science to cybersecurity and consumer protection.