Web Scraping and Data Pipelines
Collection work that turns difficult websites into usable datasets.
Built collection pipelines for cultural websites, commerce data, NFT marketplaces, and Reddit discussions. The work covered extraction, cleaning, processing, deduplication, and cloud staging for later analysis or model training.
Most model work begins before a model exists. The useful dataset has to survive changing pages, repeated records, rate limits, failed requests, and inconsistent fields.
Making collection repeatable
The pipelines used retries, rate limits, multiple workers, and deduplication so a partial failure did not throw away the whole run. Each source then had its own cleaning and processing steps.
Preparing the next step
Cleaned data was staged for analysis or model training, with the collection and transformation steps kept separate from the later experiments.
selected tools
- Python
- Selenium
- Scrapy
- BeautifulSoup
- AWS S3