Skip to content
aranyaadheu
← All projects

AI WEB EXTRACTION

deepcrawler.ai

An interactive web extraction tool that turns page content into answers to a user-defined extraction prompt, using Selenium and a local language model.

View GitHub repository →
deepcrawler.ai interface extracting quotes and authors from quotes.toscrape.com into a table
Repository screenshot: a URL and extraction prompt produce a table of quotes and authors.

At a glance

Interface
Streamlit
Local model
Llama 3.1 8B
Text chunks
Up to 6,000 characters
Repository example
Quotes extracted into a table

The problem

Web pages contain useful information mixed with navigation, scripts, and styling. Extracting just the relevant information often means writing a new parser for each task.

What I built

Built a modular Python application that captures a page with Selenium, cleans its body text with BeautifulSoup, and sends text chunks plus an extraction request to a local Ollama model through LangChain.

How it works

  1. Enter a URL and capture the page HTML using Selenium and Chrome.
  2. Remove scripts and styles, clean whitespace, and inspect the extracted text in Streamlit.
  3. Describe the information to extract and split the text into chunks of up to 6,000 characters.
  4. Process each chunk with Llama 3.1 8B through Ollama and display the combined responses.

Outcome

A Streamlit workflow for scraping, inspecting page text, and extracting requested information.

Limitations

Explore the work