Vitaliy P.
Data Engineer | ETL, SQL, PostgreSQL | RAG Pipelines & LLM Pipelines
I build RAG and LLM systems on top of business data โ and I measure whether they actually work, instead of shipping a demo that looks fine until the first real question. On a document-retrieval system I built, structure-aware chunking moved recall@5 from 0.756 to 1.000. On a text-to-SQL assistant over a 300-table warehouse schema, a semantic layer moved execution accuracy from 62.8% to 81.8% and collapsed run-to-run variance from a 13-point spread to zero. Behind that sits 10+ years of data engineering โ ETL, data warehouses and reporting for telecom, banking and FMCG โ which is why the pipelines feeding these systems keep running after handover. WHAT I BUILD - Context engineering โ deciding what actually goes into the prompt. On a text-to-SQL system, dumping the full schema cost ~22k tokens per question, roughly 18k of it irrelevant tables. Cutting that down is where both the accuracy and the bill live - RAG and semantic search over your own documents, with a retrieval evaluation you can read: recall@k, nDCG, MRR โ not "it seems to work" - LLM pipelines that stay cheap in production. My news-automation platform classifies, deduplicates, clusters and writes editor-ready drafts at 150-200 messages a day for roughly $10-15/month in inference, using cheap-model triage and prompt caching - Text-to-SQL and natural-language reporting over an existing warehouse schema - ETL and data pipelines: extraction, validation, structuring, scheduling, monitoring - PostgreSQL performance work โ slow queries, indexes, query plans - Power BI dashboards and reporting datamarts WHERE I CAME FROM 10+ years building data warehouses, ETL and BI in telecom, banking and FMCG. A datamart I built for a bank's mobile scoring model contributed roughly a 5% scoring improvement where 2% is typical. I designed a Data Vault warehouse concept with raw-layer extraction from 12 source systems. Oracle Certified Professional; PL/SQL and SQL certified. On Upwork: Top Rated, 100% Job Success, 187 hours. A current client on an analytics contract wrote: "Vitaliy is one of top players on our team." HOW I WORK Contract / B2B, remote worldwide, no relocation. Based in Almaty (UTC+5) โ I overlap with European mornings and with US East Coast mornings on request. English: full professional. I answer within 0-4 hours on working days. Stack: Python, PostgreSQL, pgvector, SQL, Claude and OpenAI APIs, embeddings (bge-m3), Docker, SQLAlchemy, Pandas, PySpark, Power BI, MS SQL Server, Oracle. If you have documents, a database or a reporting stack that an LLM is supposed to work on top of โ send me what the system is supposed to answer and what it gets wrong today, and I will tell you whether the fix is retrieval, the data underneath, or the prompt.