Carlos A.
Web Scraping | Court & Public Records | Anti-Bot | Data Pipelines
I turn hard-to-access websites and public records into clean, reliable data feeds, delivered into the tools your team already uses. My work includes 150,000+ products tracked every 2-5 hours, 250+ county record sources integrated, and 96,000+ business entities extracted and classified. I help real estate investors, proptech companies, data providers and price intelligence teams get the information they need, with extraction, updates and data quality handled as one system. REAL ESTATE DATA AT SCALE I recently completed a pipeline integrating Zillow, Wiredata and Flexmls. I am also completing a pipeline across nine Swiss real estate websites, processing 180,000+ property listings daily with extraction, updates, normalization and deduplication, including sources protected by DataDome and Akamai. A separate project underway targets 38.7 million historical Comparis listing records behind DataDome. Together, these projects cover both ongoing market monitoring and large historical backfills. COURT AND PUBLIC RECORDS My government-data work comes together under CivicMine: court filings, property records, business registries and municipal meetings, structured for research, monitoring and investment workflows. I build around the platforms behind those sources, including Tyler Odyssey, Socrata, Granicus and CivicPlus. When jurisdictions share a platform, I can reuse the extraction approach and focus on local coverage, validation and the information your team needs. The output can be ownership signals linked to properties, searchable meeting records or scheduled reports with source references. Previous work includes 5,000+ municipal meetings processed for topic detection and 100,000+ property listings deduplicated and enriched with tax and valuation data. PROTECTED SOURCES AND COMMERCIAL DATA My experience includes Cloudflare, Akamai, DataDome, PerimeterX and Imperva, across retail, automotive, marketplaces and authenticated research sources. Commercial pipelines include 150,000+ grocery products monitored every 2-5 hours across UK retailers and 125,000+ automotive parts synchronized daily. I select the extraction approach around the source, required coverage and update frequency. FROM EXTRACTION TO A RUNNING SYSTEM I handle collection, normalization, deduplication, entity resolution and delivery into PostgreSQL, Supabase, Google Sheets or your API. Where needed, I add classification and enrichment so your team receives records it can use directly. Monitoring, retries, change detection and data-quality checks make missing data and source changes visible. You receive the running system, configuration and documentation, with ongoing maintenance available. Core tools: Python, httpx, Scrapy, Playwright, Selenium and PostgreSQL. HOW WE START Send me one or two target URLs and an example of the output you need. I'll give you a free initial technical review and recommend a practical approach. Once the scope is clear, I can quote a fixed price or work hourly, depending on what suits the project.