WEB SCRAPING
Reed Scrape
A Selenium notebook that collects online-course listings across ten Reed pages and exports titles, providers, student counts, and URLs with pandas.
- Python
- Selenium
- pandas
- Jupyter
At a glance
- Pagination
- 10 pages
- Saved records
- 250
- Unique course URLs
- 250
- Dataset fields
- Title · Provider · Students · URL
The problem
Comparing course listings manually across multiple pages is repetitive. A structured dataset makes those listings easier to inspect and analyse.
What I built
Built a Selenium pagination workflow in a Jupyter notebook. It extracts listing fields using CSS selectors, collects rows in a pandas DataFrame, and exports a CSV without an extra index column.
How it works
- Open Reed online-course listing pages 1 through 10 with Chrome.
- Extract course title, provider, student-count text, and the course URL.
- Collect the records in a pandas DataFrame.
- Export courses_info.csv for further analysis.
Outcome
The committed CSV contains 250 course records with 250 unique URLs.
Limitations
- The dataset is a saved snapshot, not a continuously updated course feed.
- CSS selectors and fixed waits depend on the website's current markup and loading behaviour.
- Provider and student-count values are collected as displayed text and may be empty; further cleaning is needed for numerical analysis.