Dmitry A.
AI Data Engineer, Database Architect | SQL dbt Databricks RAG Agents
I am a dedicated Data Engineer with a strong foundation in data analytics, specializing in building production-grade data infrastructure, distributed database systems, scalable ETL pipelines, and AI-powered data applications. With 7 years of experience, I help companies: ⢠Transform messy data into automated, reliable, and scalable ETL pipelines for complex financial data that save 100+ hours monthly ⢠Architect distributed database systems: Design multi-tier data architectures for massive-scale record management ⢠Build advanced ETL pipelines with intelligent change detection to process only what's changed, reducing unnecessary processing ⢠Architect cloud-native data solutions and scale data infrastructure across multiple clouds ⢠Automate data workflows to eliminate manual processing and escape Excel Hell ⢠Build Information Retrieval (IR), semantic search, OpenSearch, and Retrieval-Augmented Generation (RAG) systems that make massive datasets, documents, and knowledge bases searchable and actionable ⢠Develop AI-powered data applications, including document intelligence, chatbot, and agent-based solutions integrated with existing business workflows ⢠Modernize legacy data systems with cloud & AI solutions ⢠Optimize data infrastructure to reduce costs and improve performance ⢠Design parallel execution frameworks: Implement isolation patterns enabling concurrent pipeline runs without conflicts ⢠Build data foundations for AI initiatives, including vector search, knowledge retrieval, document processing, and AI automation workflows I have been providing a wide range of services in the realm of data analytics and data engineering such as: ⢠ETL/ELT pipeline orchestration: Prefect, Databricks Workflows & Asset Bundles ⢠Lakehouse & distributed database architecture: Databricks/Delta, YugabyteDB (distributed PostgreSQL), Neo4j, OpenSearch ⢠Custom dbt materializations and incremental models: SCD-2, temporal tables ⢠Database optimization and storage compression ⢠Data validation, quality quarantine & reconciliation frameworks ⢠Per-run schema isolation and parallel pipeline execution ⢠Google Sheets automation and dashboard generation ⢠Excel-to-database migration and formula translation ⢠Multi-source data integration: CSV, Parquet, S3, APIs, databases ⢠Data Cleaning & Transformation at scale ⢠Operational Efficiency Analysis ⢠Financial metrics calculation systems ⢠Multi-cloud architecture design ⢠Infrastructure as Code ⢠Cloud services integration: AWS, GCP, Azure ⢠RAG & LLM data applications: grounded citations, document extraction, NL-to-SQL semantic layers ⢠Automated reporting solutions (Python-Excel integration) ⢠Data Visualization & Dashboarding While the above services encapsulate my core offerings, I am inherently adaptable and thrive on diving into new challenges and expanding my skill set. Seeking great, enthusiastic projects that will provide me with challenging, interesting work that I can learn from and contribute to. My stack: Data Engineering: ā Python ā SQL (PostgreSQL, MySQL, SQL Server, YugabyteDB, DuckDB) ā Prefect ā dbt ā PySpark ā Delta Lake ā Databricks Cloud & Infrastructure: ā AWS (EC2, S3, Glue, RDS, Lambda, EKS, DynamoDB, ECR, Bedrock) ā Google Cloud Platform (BigQuery, GKE, Bigtable, Cloud Functions) ā Azure (ADF, Synapse, AKS, Cosmos DB, Azure Functions, ACR, Text Analytics) ā Terraform ā Docker ā Kubernetes ā Prometheus Data Storages: ā RDBMS (PostgreSQL, MySQL, SQL Server, DuckDB) ā Object Storage (S3, Wasabi) ā Graph Database (Neo4j) ā Key-Value Database (Redis, DynamoDB) ā Document Database + Search Engine (OpenSearch) ETL & Data Processing: ā Pandas ā NumPy ā Selenium ā BeautifulSoup Spreadsheet Automation: ā Google Sheets API (gspread) ā Excel automation (openpyxl, xlwings) ā Automated dashboard generation Data Visualization: ā Matplotlib ā Seaborn ā Plotly ā Power BI ā Grafana Backend Development: ā FastAPI ā Flask ā RESTful APIs ā GraphQL ā Redis ā Nginx ā Gunicorn ā WebSocket AI/LLM & Agent Systems: ā OpenAI API ā Anthropic API ā Gemini API ā AWS Bedrock ā Agent Development ā AI Chatbots ā RAG ā NL-to-SQL semantic layers ā Semantic search