Purity K.
Python developer & Data engineering | ETL Pipelines,Finance,LLM,Claude
I build production data systems that run correctly after I leave β pipelines, LLM workflows, backtesting engines, and Excel/VBA automation. Not prototypes. Not one-off scripts. Former Amazon Data Engineer. Doctor of Engineering, Data Engineering, University of Florida. WHAT I'VE DELIVERED IN PRODUCTION: β Claude API & LLM workflows β document automation and agent pipelines using the Anthropic API, integrated into existing data systems. β Weekly Tech Tracker β config-driven Python pipeline processing 280,930 rows across 66 companies every week. One config change adds a new grouping. No code deploy required. β Valuation Anchor Engine β 817-day backtesting system across AMZN, GOOGL, MSFT, ORCL. HAC/Newey-West corrected, FDR-adjusted. 100% in-sample hit rate on best-anchor signals β reported honestly as in-sample, not a forward guarantee. β Macro Data Pipeline β 20+ FRED series, zero manual refresh steps, live Looker Studio dashboard. Client scaled the engagement from 30 to 40 hours/week after the first delivery. β JST Service Awards β VBA Excel automation across 14 business units, 1,000+ employees. Full refresh under 60 seconds, zero formula errors, CHRO-approved. β National ML Infrastructure β Random Forest/XGBoost ensemble (0.89 AUC-ROC), deployed across 3 national projects, adopted by an international development partner as their primary screening tool. 47 regions, 160+ staff trained. WHERE I FIT BEST: - Claude & LLM integration β API workflows, document automation, RAG pipelines, agent workflows - Data engineering & ETL β Python, SQL, Snowflake, BigQuery, dbt, config-driven pipelines - Quant finance β backtesting, statistical validation (HAC/NW, FDR), financial modeling - Excel/VBA & BI automation β Power BI, Looker Studio, Tableau WHAT SEPARATES MY WORK: Config-driven by design β output updates when an input changes, no developer required after handover. QA runs automatically on every execution, not as an afterthought. I separate in-sample results from forward-tested ones and say so plainly, before a client finds out the hard way. Python Β· SQL Β· Anthropic Claude API Β· LLM Prompt Engineering Β· Snowflake Β· BigQuery Β· dbt Β· Google Apps Script Β· Looker Studio Β· Power BI Β· pandas/NumPy/scipy Β· statsmodels Β· scikit-learn Β· XGBoost Β· VBA Excel Doctor of Engineering, Data Engineering, University of Florida. Previously Amazon (data engineering, ETL, Power BI). Trilingual: English, Swahili, Spanish. Send me what you're working with β I'll give you a straight answer on fit, timeline, and cost before you spend a single connect.