Haikal H.
Large-Scale Web Scraping & Ingestion Pipelines | Python
I build and maintain large-scale Python web scraping and data ingestion pipelines, including a production system covering approximately 2,900 sources and more than 226,000 articles per week. I help teams turn unreliable public-data collection into monitored, maintainable systems with measurable source coverage, data-quality validation, failure alerting, and clean downstream delivery. My production experience includes around 100 custom parser adaptations and crawler workloads running across an eight-node environment. I build Python systems that collect data from public APIs, sitemaps, feeds, and dynamic websites, then validate, normalize, deduplicate, and deliver it as clean CSV, JSON, database records, or APIs. I can help with: ⢠Recurring web and API data collection ⢠Scrapy, Playwright, and requests-based collectors ⢠Multi-source data ingestion pipelines ⢠CSV, JSON, database, and API delivery ⢠Deduplication and data-quality validation ⢠Scheduling, monitoring, and failure alerting ⢠Docker deployment and long-term maintenance My focus is not simply making a scraper run once. I build systems where source coverage is measurable, failures are visible, and downstream data remains safe to use. I prioritize official APIs and public endpoints, respectful rate limits, incremental collection, reusable adapters, and maintainable implementations instead of fragile one-off scripts.