Skip to content
Divyesh-1729Public

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 

Repository files navigation

Reddit Post Classifier A machine learning project that automatically classifies Reddit posts into relevant topics using natural language processing techniques. 🎯 Overview This project implements a text classification system that analyzes Reddit post content and categorizes them into predefined topics. The classifier uses various machine learning algorithms and NLP techniques to achieve accurate topic prediction. ✨ Features

Text Preprocessing Pipeline: Tokenization, stopword removal, and TF-IDF vectorization Multiple ML Models: Comparison between Logistic Regression, Random Forest, and Naïve Bayes Performance Evaluation: Comprehensive metrics including accuracy, precision, recall, and F1-score Scalable Architecture: Designed to handle large datasets efficiently

🛠️ Technologies Used

Python: Core programming language scikit-learn: Machine learning algorithms and evaluation metrics Pandas: Data manipulation and analysis NLTK: Natural language processing toolkit Matplotlib: Data visualization and model performance plots

📋 Prerequisites bashPython 3.7+ pip (Python package installer) 🚀 Installation

Clone the repository:

bashgit clone https://github.com/Divyesh-1729/reddit-post-classifier.git cd reddit-post-classifier

Install required packages:

bashpip install -r requirements.txt

Download NLTK data:

pythonimport nltk nltk.download('stopwords') nltk.download('punkt') 💻 Usage

Data Preparation:

pythonpython data_preprocessing.py

Train Models:

pythonpython train_models.py

Evaluate Performance:

pythonpython evaluate_models.py

Make Predictions:

pythonpython predict.py --text "Your Reddit post text here" 📊 Model Performance ModelAccuracyPrecisionRecallF1-ScoreLogistic Regression85.2%84.7%85.1%84.9%Random Forest83.8%83.2%84.1%83.6%Naïve Bayes82.1%81.9%82.3%82.1% 📁 Project Structure reddit-post-classifier/ │ ├── data/ │ ├── raw/ # Raw Reddit data │ └── processed/ # Preprocessed data │ ├── src/ │ ├── preprocessing.py # Text preprocessing functions │ ├── models.py # ML model implementations │ ├── evaluation.py # Model evaluation metrics │ └── utils.py # Utility functions │ ├── notebooks/ │ └── exploratory_analysis.ipynb │ ├── requirements.txt ├── README.md └── main.py 🔍 Key Features Explained Text Preprocessing

Tokenization: Splits text into individual words/tokens Stopword Removal: Eliminates common words that don't add meaning TF-IDF Vectorization: Converts text to numerical features based on term frequency

Model Comparison The project implements three different algorithms:

Logistic Regression: Linear model with high interpretability Random Forest: Ensemble method with feature importance ranking Naïve Bayes: Probabilistic classifier optimized for text data

🤝 Contributing

Fork the repository Create a feature branch (git checkout -b feature/new-feature) Commit your changes (git commit -am 'Add new feature') Push to the branch (git push origin feature/new-feature) Create a Pull Request

📄 License This project is licensed under the MIT License - see the LICENSE file for details. 👨‍💻 Author Divyesh Puranik

GitHub: @Divyesh-1729 LinkedIn: Divyesh Puranik Email: divyeshpuranik@gmail.com

🙏 Acknowledgments

Reddit API for providing access to post data scikit-learn community for excellent ML tools NLTK team for comprehensive NLP resources

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors