Reddit Post Classifier A machine learning project that automatically classifies Reddit posts into relevant topics using natural language processing techniques. 🎯 Overview This project implements a text classification system that analyzes Reddit post content and categorizes them into predefined topics. The classifier uses various machine learning algorithms and NLP techniques to achieve accurate topic prediction. ✨ Features
Text Preprocessing Pipeline: Tokenization, stopword removal, and TF-IDF vectorization Multiple ML Models: Comparison between Logistic Regression, Random Forest, and Naïve Bayes Performance Evaluation: Comprehensive metrics including accuracy, precision, recall, and F1-score Scalable Architecture: Designed to handle large datasets efficiently
🛠️ Technologies Used
Python: Core programming language scikit-learn: Machine learning algorithms and evaluation metrics Pandas: Data manipulation and analysis NLTK: Natural language processing toolkit Matplotlib: Data visualization and model performance plots
📋 Prerequisites bashPython 3.7+ pip (Python package installer) 🚀 Installation
Clone the repository:
bashgit clone https://github.com/Divyesh-1729/reddit-post-classifier.git cd reddit-post-classifier
Install required packages:
bashpip install -r requirements.txt
Download NLTK data:
pythonimport nltk nltk.download('stopwords') nltk.download('punkt') 💻 Usage
Data Preparation:
pythonpython data_preprocessing.py
Train Models:
pythonpython train_models.py
Evaluate Performance:
pythonpython evaluate_models.py
Make Predictions:
pythonpython predict.py --text "Your Reddit post text here" 📊 Model Performance ModelAccuracyPrecisionRecallF1-ScoreLogistic Regression85.2%84.7%85.1%84.9%Random Forest83.8%83.2%84.1%83.6%Naïve Bayes82.1%81.9%82.3%82.1% 📁 Project Structure reddit-post-classifier/ │ ├── data/ │ ├── raw/ # Raw Reddit data │ └── processed/ # Preprocessed data │ ├── src/ │ ├── preprocessing.py # Text preprocessing functions │ ├── models.py # ML model implementations │ ├── evaluation.py # Model evaluation metrics │ └── utils.py # Utility functions │ ├── notebooks/ │ └── exploratory_analysis.ipynb │ ├── requirements.txt ├── README.md └── main.py 🔍 Key Features Explained Text Preprocessing
Tokenization: Splits text into individual words/tokens Stopword Removal: Eliminates common words that don't add meaning TF-IDF Vectorization: Converts text to numerical features based on term frequency
Model Comparison The project implements three different algorithms:
Logistic Regression: Linear model with high interpretability Random Forest: Ensemble method with feature importance ranking Naïve Bayes: Probabilistic classifier optimized for text data
🤝 Contributing
Fork the repository Create a feature branch (git checkout -b feature/new-feature) Commit your changes (git commit -am 'Add new feature') Push to the branch (git push origin feature/new-feature) Create a Pull Request
📄 License This project is licensed under the MIT License - see the LICENSE file for details. 👨💻 Author Divyesh Puranik
GitHub: @Divyesh-1729 LinkedIn: Divyesh Puranik Email: divyeshpuranik@gmail.com
🙏 Acknowledgments
Reddit API for providing access to post data scikit-learn community for excellent ML tools NLTK team for comprehensive NLP resources