FNU S.
Python Data Engineer | Anti-Bot Web Scraping, Cleaning, Automation
You have data that is messy, scattered, or stuck behind a site that does not want to give it up. I get it out, clean it, match and dedupe it, and hand it back as a spreadsheet, a database, or an API you can call. I am a Top Rated Python data engineer. Most of my work is public and live. The demos and the source are in the Portfolio section of this profile. What I do: - Scraping protected, anti-bot sites. This is the part I am known for. Sites behind Cloudflare, Incapsula, CAPTCHA walls, rate limits, and the usual bot-detection and WAF defenses. The tooling is Playwright with stealth, proxy rotation, and backoff and retry. Selenium when a page is JavaScript-heavy or login-walled, BeautifulSoup when the HTML is plain. Pagination handled, clean CSV or Excel out. One-time pull, or a scheduled job with change detection so you only see new or changed rows. - Cleaning, fuzzy matching, dedupe. Normalize, validate, then RapidFuzz multi-scorer matching to fold messy records from different sources into one clean dataset. This is the work clients keep coming back for. - Automation and pipelines. Scheduled jobs and ETL into PostgreSQL, Redis in front of it, data-quality and anomaly checks, FastAPI or Flask backends, Docker to deploy. Scrape, clean, match, load, deliver. Hands off after the first run. - API integrations and small apps. FastAPI plus React, third-party APIs (payments, translation, news and RSS feeds), JWT auth, webhook handling. The Portfolio section has the live ones. A data-quality and anomaly-detection framework I put in front of a pipeline load. A Dockerized FastAPI and React credit-risk system behind a JWT-auth API. A news-sentiment pipeline that ingests, cleans, analyzes and serves, live. A card-authorization API. More than that, each with a running URL and its source on my linked GitHub. The reviews on this profile say it better than I can. Clients hire me for scraping, automation and data work, and they come back. What comes up again and again: clear communication, work that arrives clean and works, and not needing to be chased. Read them in their own words. How I work: tight scope, plain language. Before you hire I will tell you how I would build it, where the real risk is (is the site JavaScript-rendered, is it behind a login, what bot detection does it run, how much data), and what you get for the money. If a site is genuinely not worth scraping I will say so upfront instead of burning your budget proving it. For scraping and pipeline jobs I can usually send a small sample first, a few rows or a profiling pass, so you see the actual output before you commit a cent. Stack: Python 3.10+, Playwright, Selenium, BeautifulSoup, RapidFuzz, pandas, FastAPI, Flask, PostgreSQL, Redis, Docker, React, TypeScript, scikit-learn, REST APIs, scheduled jobs and ETL. AWS Certified Data Engineer Associate and AWS Certified Cloud Practitioner. Send me the source (the site, the file, the API, or the messy dataset) and what you want out of it. I will come back with how I would build it, a fixed price, and a timeline. If it is a scrape or a cleaning job I can send a sample of the real output first, so you can judge it before you hire.