Skip to content

Add text classification model to detect spam messages - #321

Closed
student-smritipandey wants to merge 3 commits into
code-a2z:mainfrom
student-smritipandey:spam-detector-model
Closed

student-smritipandey wants to merge 3 commits into
code-a2z:mainfrom
student-smritipandey:spam-detector-model

Conversation

@student-smritipandey

@student-smritipandey student-smritipandey commented Aug 17, 2025 •

Copy link
Copy Markdown

This project is a text classification model that detects whether a given SMS/text message is Spam or Ham (Not Spam) using Deep Learning. The model was trained on a labeled dataset of SMS messages and deployed with a user-friendly Streamlit web application.

🔹 Project Workflow

Dataset used: https://www.kaggle.com/datasets/uciml/sms-spam-collection-dataset

Data Preprocessing

Cleaned and tokenized SMS messages

Converted text to sequences using Tokenizer

Applied padding to maintain equal input length (MAX_LEN = 100)

Model Architecture

Built using TensorFlow/Keras

Embedding layer for word representation

LSTM / Dense layers for sequential learning

Final sigmoid output layer for binary classification (Spam vs Ham)

Training & Evaluation

Optimizer: Adam

Loss function: Binary Crossentropy

Metrics: Accuracy

Achieved high classification performance on test data

Deployment

Model saved as spam_classifier.h5 / spam_classifier.keras

Tokenizer saved as tokenizer.pkl

Deployed using Streamlit for real-time predictions

🔹 Features

✅ Detects spam messages with high accuracy
✅ Returns prediction confidence score
✅ Interactive web UI using Streamlit
✅ Supports any user-entered SMS/text message

🔹 Example Predictions

Input: "Congratulations! You have won a $500 gift voucher. Click the link to claim."
→ 🚨 Spam detected! (Confidence: 99%)

Input: "Hey, are we still meeting tomorrow at 5?"
→ ✅ Ham (Not Spam) (Confidence: 100%)

🔹 Technologies Used

Python

TensorFlow / Keras

NLTK / Text Preprocessing

Streamlit

Pickle

This project is a text classification model that detects whether a given SMS/text message is Spam or Ham (Not Spam) using Deep Learning. The model was trained on a labeled dataset of SMS messages and deployed with a user-friendly Streamlit web application.

🔹 Project Workflow

Dataset used: https://www.kaggle.com/datasets/uciml/sms-spam-collection-dataset

Data Preprocessing

Cleaned and tokenized SMS messages

Converted text to sequences using Tokenizer

Applied padding to maintain equal input length (MAX_LEN = 100)

Model Architecture

Built using TensorFlow/Keras

Embedding layer for word representation

LSTM / Dense layers for sequential learning

Final sigmoid output layer for binary classification (Spam vs Ham)

Training & Evaluation

Optimizer: Adam

Loss function: Binary Crossentropy

Metrics: Accuracy

Achieved high classification performance on test data

Deployment

Model saved as spam_classifier.h5 / spam_classifier.keras

Tokenizer saved as tokenizer.pkl

Deployed using Streamlit for real-time predictions

🔹 Features

✅ Detects spam messages with high accuracy
✅ Returns prediction confidence score
✅ Interactive web UI using Streamlit
✅ Supports any user-entered SMS/text message

🔹 Example Predictions

Input: "Congratulations! You have won a $500 gift voucher. Click the link to claim."
→ 🚨 Spam detected! (Confidence: 99%)

Input: "Hey, are we still meeting tomorrow at 5?"
→ ✅ Ham (Not Spam) (Confidence: 100%)

🔹 Technologies Used

Python

TensorFlow / Keras

NLTK / Text Preprocessing

Streamlit

Pickle
@github-actions

Copy link
Copy Markdown

Thank you for submitting your pull request! We'll review it as soon as possible. For further communication, join our discord server https://discord.gg/tSqtvHUJzE.

@Avdhesh-Varshney

Copy link
Copy Markdown
Collaborator

@student-smritipandey Have you gone through the PR review criteria in README.md file?
Plz gone through all the guidelines and 1 more extra step is to review the movieReccomendationModel.py file

Remove unnecessary files and add the spam_detector.py
@student-smritipandey

Copy link
Copy Markdown
Author

@Avdhesh-Varshney I have committed changes, review it

@Avdhesh-Varshney

Copy link
Copy Markdown
Collaborator

Hi @student-smritipandey Right now correct! But have you consider the case where will the app find spam_detector.h5 model file?

Can we connect on discord? So I will tell you what needs to be change next.
Or carefully gone through movieReccomendationModel.py file and guidelines and keep your model in ChatBot folder.

Don't create any new folder.

@student-smritipandey

Copy link
Copy Markdown
Author

@Avdhesh-Varshney yeah! give me discord link

@Avdhesh-Varshney

Copy link
Copy Markdown
Collaborator

Thank you for submitting your pull request! We'll review it as soon as possible. For further communication, join our discord server https://discord.gg/tSqtvHUJzE.

@student-smritipandey Join server and take the respective project channel OS - Jarvis role from get-roles channel

@student-smritipandey

Copy link
Copy Markdown
Author

@Avdhesh-Varshney I have uploaded the model my my kaggle profile please review it https://www.kaggle.com/code/smritipandey02/spam-detection

@Avdhesh-Varshney

Copy link
Copy Markdown
Collaborator

Nice! Now you have to load your model in Jarvis using helper functions by providing the fields with your kaggle username and notebook name as in movieRecommendationModel.py file

Also update the directory as mentioned earlier!

@student-smritipandey

Copy link
Copy Markdown
Author

@Avdhesh-Varshney updated the file

@Avdhesh-Varshney Avdhesh-Varshney added the wontfix ❌ This will not be worked on label Oct 5, 2025
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

wontfix ❌ This will not be worked on

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants