Skip to content
aranyaadheu
← All projects

WEB SCRAPING

Reed Scrape

A Selenium notebook that collects online-course listings across ten Reed pages and exports titles, providers, student counts, and URLs with pandas.

View GitHub repository →

At a glance

Pagination
10 pages
Saved records
250
Unique course URLs
250
Dataset fields
Title · Provider · Students · URL

The problem

Comparing course listings manually across multiple pages is repetitive. A structured dataset makes those listings easier to inspect and analyse.

What I built

Built a Selenium pagination workflow in a Jupyter notebook. It extracts listing fields using CSS selectors, collects rows in a pandas DataFrame, and exports a CSV without an extra index column.

How it works

  1. Open Reed online-course listing pages 1 through 10 with Chrome.
  2. Extract course title, provider, student-count text, and the course URL.
  3. Collect the records in a pandas DataFrame.
  4. Export courses_info.csv for further analysis.

Outcome

The committed CSV contains 250 course records with 250 unique URLs.

Limitations

Explore the work