Back to LeaderboardLast snapshot: Sep 9
Dmitry A.
Badge

Dmitry A.

AI Data Engineer, Database Architect | SQL dbt Databricks RAG Agents

I am a dedicated Data Engineer with a strong foundation in data analytics, specializing in building production-grade data infrastructure, distributed database systems, scalable ETL pipelines, and AI-powered data applications. With 7 years of experience, I help companies: • Transform messy data into automated, reliable, and scalable ETL pipelines for complex financial data that save 100+ hours monthly • Architect distributed database systems: Design multi-tier data architectures for massive-scale record management • Build advanced ETL pipelines with intelligent change detection to process only what's changed, reducing unnecessary processing • Architect cloud-native data solutions and scale data infrastructure across multiple clouds • Automate data workflows to eliminate manual processing and escape Excel Hell • Build Information Retrieval (IR), semantic search, OpenSearch, and Retrieval-Augmented Generation (RAG) systems that make massive datasets, documents, and knowledge bases searchable and actionable • Develop AI-powered data applications, including document intelligence, chatbot, and agent-based solutions integrated with existing business workflows • Modernize legacy data systems with cloud & AI solutions • Optimize data infrastructure to reduce costs and improve performance • Design parallel execution frameworks: Implement isolation patterns enabling concurrent pipeline runs without conflicts • Build data foundations for AI initiatives, including vector search, knowledge retrieval, document processing, and AI automation workflows I have been providing a wide range of services in the realm of data analytics and data engineering such as: • ETL/ELT pipeline orchestration: Prefect, Databricks Workflows & Asset Bundles • Lakehouse & distributed database architecture: Databricks/Delta, YugabyteDB (distributed PostgreSQL), Neo4j, OpenSearch • Custom dbt materializations and incremental models: SCD-2, temporal tables • Database optimization and storage compression • Data validation, quality quarantine & reconciliation frameworks • Per-run schema isolation and parallel pipeline execution • Google Sheets automation and dashboard generation • Excel-to-database migration and formula translation • Multi-source data integration: CSV, Parquet, S3, APIs, databases • Data Cleaning & Transformation at scale • Operational Efficiency Analysis • Financial metrics calculation systems • Multi-cloud architecture design • Infrastructure as Code • Cloud services integration: AWS, GCP, Azure • RAG & LLM data applications: grounded citations, document extraction, NL-to-SQL semantic layers • Automated reporting solutions (Python-Excel integration) • Data Visualization & Dashboarding While the above services encapsulate my core offerings, I am inherently adaptable and thrive on diving into new challenges and expanding my skill set. Seeking great, enthusiastic projects that will provide me with challenging, interesting work that I can learn from and contribute to. My stack: Data Engineering: āœ… Python āœ… SQL (PostgreSQL, MySQL, SQL Server, YugabyteDB, DuckDB) āœ… Prefect āœ… dbt āœ… PySpark āœ… Delta Lake āœ… Databricks Cloud & Infrastructure: āœ… AWS (EC2, S3, Glue, RDS, Lambda, EKS, DynamoDB, ECR, Bedrock) āœ… Google Cloud Platform (BigQuery, GKE, Bigtable, Cloud Functions) āœ… Azure (ADF, Synapse, AKS, Cosmos DB, Azure Functions, ACR, Text Analytics) āœ… Terraform āœ… Docker āœ… Kubernetes āœ… Prometheus Data Storages: āœ… RDBMS (PostgreSQL, MySQL, SQL Server, DuckDB) āœ… Object Storage (S3, Wasabi) āœ… Graph Database (Neo4j) āœ… Key-Value Database (Redis, DynamoDB) āœ… Document Database + Search Engine (OpenSearch) ETL & Data Processing: āœ… Pandas āœ… NumPy āœ… Selenium āœ… BeautifulSoup Spreadsheet Automation: āœ… Google Sheets API (gspread) āœ… Excel automation (openpyxl, xlwings) āœ… Automated dashboard generation Data Visualization: āœ… Matplotlib āœ… Seaborn āœ… Plotly āœ… Power BI āœ… Grafana Backend Development: āœ… FastAPI āœ… Flask āœ… RESTful APIs āœ… GraphQL āœ… Redis āœ… Nginx āœ… Gunicorn āœ… WebSocket AI/LLM & Agent Systems: āœ… OpenAI API āœ… Anthropic API āœ… Gemini API āœ… AWS Bedrock āœ… Agent Development āœ… AI Chatbots āœ… RAG āœ… NL-to-SQL semantic layers āœ… Semantic search

šŸ‡°šŸ‡æAlmaty, Kazakhstan$45/hr100% JSS5.00 (26)Since Jan 2018
MRR
$3,819
šŸŒ#8,696/363K
šŸ‡°šŸ‡æ#29/668
Recent Earnings
$22,913
šŸŒ#8,696/363K
šŸ‡°šŸ‡æ#29/668
Total Earnings
$176K
šŸŒ#17.3K/363K
šŸ‡°šŸ‡æ#74/668
Avg. Per Project
$4,518
šŸŒ#35.6K/363K
šŸ‡°šŸ‡æ#101/668
Recent Projects
4
1f3h
šŸŒ#72.0K/363K
šŸ‡°šŸ‡æ#146/668
Total Projects
39
26f21h
šŸŒ#72.0K/363K
šŸ‡°šŸ‡æ#146/668
MRR Performance Over Time
$50k$25k$0
6 mo ago3 mo agoNow
Coming SoonGathering historical data
World Skill Rankings
of 658
#5
RussianTop 0.6%
#14Database Optimization
#15Object-Oriented Programming
#66Data Modeling
#91Data Extraction
#91Data Mining
#96ETL Pipeline
#98Data Engineering
#140Data Science
#228Microsoft Azure
#344Data Analysis
#387Machine Learning
#534SQL
#1,236Python
šŸ‡°šŸ‡æKazakhstan Skill Rankings
of 8
#1
Data ScienceTop <0.01%
#1Russian
#1Data Analysis
#1Data Extraction
#1Data Engineering
#1Data Modeling
#1Database Optimization
#1ETL Pipeline
#1Microsoft Azure
#1Object-Oriented Programming
#2Machine Learning
#2Data Mining
#4SQL
#10Python
Dmitry A. — Top 4% in Kazakhstan | UpworkMRR