Vinicius d.
AI & Data Engineer | LLM Apps, Pipelines, MLOps | AWS, GCP, Azure
I build data platforms and AI systems that run in production โ pipelines, RAG and agent applications, evaluation and monitoring, and the infrastructure underneath. 9+ years shipping, on all three major clouds. Multi-cloud, for real: I've built and operated production platforms on AWS, GCP, and Azure โ not one cloud plus certifications in the other two. That includes BigQuery at 4 PB/day, Kubernetes (EKS/GKE/AKS) for GPU workloads, and Terraform-managed infrastructure across all three. B2B SaaS and regulated industries: I currently work on an enterprise AI observability platform, and most of my delivery experience is with banking, insurance, and telecom clients โ environments where audit trails, PII handling, data residency, and access controls aren't optional. Familiar with EU AI Act, ISO 42001, NIST AI RMF, AIUC, and LGPD/GDPR requirements as they land on engineering teams. Recent work: - Multi-cluster GPU job orchestration on Azure Kubernetes (AKS Fleet Manager, MultiKueue) for an autonomous trucking company - Airflow + MLflow platform serving multiple ML teams โ DAG design, version migrations, CI/CD, cost control - Data platform processing 4 PB/day on BigQuery โ 50% cloud cost reduction, 85% data quality improvement, team of 25 - Enterprise LLM evaluation and observability: tracing, guardrails, LLM-as-judge, regression testing - ML research on passive sonar with the Brazilian Navy โ published internationally, prototype deployed on a submarine Data engineering - Batch and streaming pipelines, warehouse and lakehouse design, dbt modeling, data quality and observability, migrations, cost optimization. PII handling, lineage, and audit-ready pipelines for regulated workloads. Airflow, Spark, Kafka, dbt, BigQuery, Databricks, Snowflake and Redshift. AI engineering - RAG systems, LLM agents, evaluation pipelines, guardrails, tracing and monitoring, model serving. LangChain/LangGraph, vector databases, OpenAI/Anthropic/Bedrock/Vertex AI. ML infrastructure - Training and inference on Kubernetes, GPU scheduling, MLflow, feature pipelines, CI/CD for ML, Terraform. EKS, GKE, AKS. I'm useful whether you need a pipeline built from scratch, a broken one diagnosed, or a proof-of-concept turned into something that won't fall over at scale. Most engagements start small and run long โ 14 contracts, 2,500+ hours, $200K+ earned on Upwork. Stack: Python, SQL, PyTorch, FastAPI, Airflow, dbt, Spark, BigQuery, Databricks, Snowflake, MLflow, Kubernetes, Docker, Terraform, AWS, GCP, Azure. Tell me what you're building or what's broken. If I'm not the right fit, I'll say so.