This repository contains Python code showcasing the evolution of a web scraping project, starting from basic data extraction using BeautifulSoup to building a versatile web crawler with automation capabilities using Selenium. The project gradually progresses to data storage in MongoDB and concludes with an automated data filler for social media accounts.
-
Introduction As a beginner in web scraping, this project documents my learning journey and the progression of techniques used to extract and manipulate data from various websites.
-
Listing and Installing Dependencies a. Python 3.x b. BeautifulSoup c. Selenium d. MongoDB e. pymongo
2.1 Installing the Dependencies
pip install beautifulsoup4 selenium pymongo-
Project Structure The project is organized into different stages, each representing a milestone in the learning process. The structure is as follows: ├── bs4_extraction ├── selenium_scraping ├── mongodb_integration ├── versatile_web_crawler ├── social_media_automation
-
Web Scraping with BeautifulSoup In the initial stage, the project focuses on extracting data from a specific website (Coingecko) using the BeautifulSoup library. The extracted data is displayed using arrays to make it more readable.
-
Advancing to Selenium Recognizing the limitations of BeautifulSoup, the project transitions to using Selenium for web scraping. It includes examples of scraping data from pages that require loading more data dynamically. The code demonstrates proficiency in handling dynamic content.
-
Data Storage in MongoDB Building on the knowledge gained, the project introduces MongoDB integration. Data scraped using Selenium is now stored in a MongoDB database. The code includes URI connection and usage of the pymongo library.
-
Versatile Web Crawler The web crawler developed in this stage is capable of scraping data from any given website. It simplifies the extracted data and stores it in MongoDB. The crawler is designed to update existing data if newer information is found, while preserving the old data.
-
Automation for Social Media The final stage involves creating an automated data filler for social media accounts. The code assists in creating accounts using information provided within the script. It showcases automation skills beyond web scraping.
To use Selenium for web scraping, you need to download the Chrome WebDriver. Follow the steps below to add Chrome WebDriver to your project:
-
Check Your Chrome Browser Version:
- Open Chrome.
- Click on the three dots in the top-right corner.
- Go to "Help" > "About Google Chrome."
- Note the version number.
-
Download Chrome WebDriver:
- Visit the ChromeDriver Downloads page.
- Download the version of ChromeDriver that matches your Chrome browser version.
-
Extract the Chrome WebDriver:
- Extract the downloaded ZIP file to get the
chromedriver.exefile.
- Extract the downloaded ZIP file to get the
-
Place the Chrome WebDriver in Your Project Directory:
- Copy the
chromedriver.exefile to the directory where your Selenium script is located.
- Copy the
Now, you are ready to use Selenium with Chrome WebDriver in your project.
# Example Selenium script using Chrome WebDriver
from selenium import webdriver
# Specify the path to the chromedriver.exe file
chrome_path = './chromedriver.exe'
# Create a new instance of the Chrome driver
driver = webdriver.Chrome(executable_path=chrome_path)Feel free to explore for detailed code and examples associated with each stage of the project. Happy coding!