Weikang Z.
Data Scraping Expert | Data Pipeline | Data Analysis
Encountering hardships when scraping data from platforms with tough anti-scraping measures or unstructured websites? Even after getting the data, you might be puzzled about processing it, like converting video and audio info into usable data. And setting up a storage structure for large-scale data, along with the inefficiencies in data cleaning and analysis, can be daunting. Don't worry, Rely on my 6+ years of experience in the data industry, and I will effectively help you solve these problems. 📊 I hold a master's degree in Computer Science and have over six years of valuable tech industry experience. I worked at VIVO Communication Technology Co., Ltd, a top 50 Chinese Tech firm for six years, focusing on massive data scraping, modeling, warehousing, and analysis. This has given me deep expertise and rich practical knowledge in various data processing aspects. Work Experience and Skills 👉Data Acquisition Against Anti-scraping: I have processed multiple data websites that are irregular in structure and have strict anti-scraping measures, and even have mechanisms that are ready to identify and block scraping attempts. I am particularly proficient in analyzing the anti-scraping logic of platforms such as TikTok, Reddit, and Instagram with strict anti-scraping mechanisms. Tools like Selenium, Scrapy, and Playwright help me legally get valuable data. My customized scraping strategies break through barriers, providing a solid data base for analysis. 👉Efficient Batch Transcription: With tools such as Gemini, GPT and their APIs, I batch-transcribe videos, audios, and texts based on well-designed prompts, extracting needed data accurately. In TikTok and Douyin projects, I quickly transcribed much audio, supporting social trend analysis. 👉Professional ETL Data Processing: In ETL, I use Python's Pandas and NumPy to clean and preprocess scraped and transcribed data. My efficient process ensures data quality, preparing for in-depth analysis and improving result accuracy. 👉Machine Learning Data Analysis: Familiar with algorithms like LightGBM, XGBoost, I analyze processed data, uncovering hidden value and supporting advertising strategy optimization for businesses. 👉Data Storage Structure Building: With rich data engineering experience, I choose suitable databases (MySQL, PostgreSQL, MongoDB and Hive) based on data scale and needs. I create efficient, stable, scalable storage structures, ensuring data security and improving read/write performance for easy retrieval and analysis. 👉Process Automation: Proficient in Shell scripts, Python scripts, I automate the whole data process. Git for version control ensures code traceability and teamwork efficiency. Docker for environment deployment ensures process consistency. Automation in scraping, transcription, cleaning, and analysis improves efficiency, reduces errors, and enhances accuracy. 👉Exception Alarm Mechanism: I designed an intelligent alarm system with real-time monitoring, threshold detection, and early warning. It spots data anomalies by comparing indicators with set thresholds. When an exception occurs, it sends detailed info via email, ensuring data integrity. Tools and Technologies I Master ✨ Python, SQL, VBA, Shell Script, Airflow,AWS,git; ✨ LightGBM, XGBoost, KNN, Matplotlib, Pandas, NumPy; ✨ BeautifulSoup, Selenium, Scrapy, Playwright; ✨ Tableau, Power BI; 🤝 Customer Recognition and Long-term Cooperation I've worked on large data scraping projects on Upwork. Despite their complexity, I achieved great results, winning customer trust. Over 50% of my business is long-term cooperation, as customers value my ability, attitude, and the value I bring. I respond quickly to technical issues and data needs, showing I can create value and meet changing demands. 👋I have practical data processing and analysis experience and can solve data problems. I look forward to working with you to explore data value and drive business growth.