Back to LeaderboardLast snapshot: Sep 8
Vitaliy P.
Badge

Vitaliy P.

Data Engineer | ETL, SQL, PostgreSQL | RAG Pipelines & LLM Pipelines

I build RAG and LLM systems on top of business data โ€” and I measure whether they actually work, instead of shipping a demo that looks fine until the first real question. On a document-retrieval system I built, structure-aware chunking moved recall@5 from 0.756 to 1.000. On a text-to-SQL assistant over a 300-table warehouse schema, a semantic layer moved execution accuracy from 62.8% to 81.8% and collapsed run-to-run variance from a 13-point spread to zero. Behind that sits 10+ years of data engineering โ€” ETL, data warehouses and reporting for telecom, banking and FMCG โ€” which is why the pipelines feeding these systems keep running after handover. WHAT I BUILD - Context engineering โ€” deciding what actually goes into the prompt. On a text-to-SQL system, dumping the full schema cost ~22k tokens per question, roughly 18k of it irrelevant tables. Cutting that down is where both the accuracy and the bill live - RAG and semantic search over your own documents, with a retrieval evaluation you can read: recall@k, nDCG, MRR โ€” not "it seems to work" - LLM pipelines that stay cheap in production. My news-automation platform classifies, deduplicates, clusters and writes editor-ready drafts at 150-200 messages a day for roughly $10-15/month in inference, using cheap-model triage and prompt caching - Text-to-SQL and natural-language reporting over an existing warehouse schema - ETL and data pipelines: extraction, validation, structuring, scheduling, monitoring - PostgreSQL performance work โ€” slow queries, indexes, query plans - Power BI dashboards and reporting datamarts WHERE I CAME FROM 10+ years building data warehouses, ETL and BI in telecom, banking and FMCG. A datamart I built for a bank's mobile scoring model contributed roughly a 5% scoring improvement where 2% is typical. I designed a Data Vault warehouse concept with raw-layer extraction from 12 source systems. Oracle Certified Professional; PL/SQL and SQL certified. On Upwork: Top Rated, 100% Job Success, 187 hours. A current client on an analytics contract wrote: "Vitaliy is one of top players on our team." HOW I WORK Contract / B2B, remote worldwide, no relocation. Based in Almaty (UTC+5) โ€” I overlap with European mornings and with US East Coast mornings on request. English: full professional. I answer within 0-4 hours on working days. Stack: Python, PostgreSQL, pgvector, SQL, Claude and OpenAI APIs, embeddings (bge-m3), Docker, SQLAlchemy, Pandas, PySpark, Power BI, MS SQL Server, Oracle. If you have documents, a database or a reporting stack that an LLM is supposed to work on top of โ€” send me what the system is supposed to answer and what it gets wrong today, and I will tell you whether the fix is retrieval, the data underneath, or the prompt.

๐Ÿ‡ฐ๐Ÿ‡ฟAlmaty, Kazakhstan$37/hr100% JSS4.99 (6)Since Jul 2020
MRR
$301
๐ŸŒ#92.6K/363K
๐Ÿ‡ฐ๐Ÿ‡ฟ#157/668
Recent Earnings
$1,807
๐ŸŒ#92.6K/363K
๐Ÿ‡ฐ๐Ÿ‡ฟ#157/668
Total Earnings
$5,598
๐ŸŒ#227K/363K
๐Ÿ‡ฐ๐Ÿ‡ฟ#439/668
Avg. Per Project
$622
๐ŸŒ#150K/363K
๐Ÿ‡ฐ๐Ÿ‡ฟ#307/668
Recent Projects
2
0f2h
๐ŸŒ#193K/363K
๐Ÿ‡ฐ๐Ÿ‡ฟ#383/668
Total Projects
9
5f4h
๐ŸŒ#193K/363K
๐Ÿ‡ฐ๐Ÿ‡ฟ#383/668
MRR Performance Over Time
$50k$25k$0
6 mo ago3 mo agoNow
Coming SoonGathering historical data
World Skill Rankings
of 173
#56
Database OptimizationTop 32%
#127Data Warehousing
#245Microsoft Power BI Data Visualization
#270Data Modeling
#284ChatGPT API Integration
#340ETL Pipeline
#356Vector Database
#409Data Engineering
#674Large Language Model
#728Prompt Engineering
#805Chatbot Development
#876Web Scraping
#972Retrieval Augmented Generation
#1,202Data Extraction
#1,986AI Development
#2,562SQL
#2,928PostgreSQL
#5,059API Integration
#6,778Python
๐Ÿ‡ฐ๐Ÿ‡ฟKazakhstan Skill Rankings
of 4
#1
Large Language ModelTop <0.01%
#1Data Warehousing
#1Microsoft Power BI Data Visualization
#2Vector Database
#2ChatGPT API Integration
#2Database Optimization
#2ETL Pipeline
#2Data Engineering
#2Data Modeling
#3Retrieval Augmented Generation
#3Prompt Engineering
#3Chatbot Development
#3Data Extraction
#4AI Development
#4Web Scraping
#14SQL
#15PostgreSQL
#15API Integration
#27Python
Vitaliy P. โ€” Top 23% in Kazakhstan | UpworkMRR