Mehrshad L.
AI/LLM Observability Expert | Grafana, MLOps, Cost Monitoring
AI/LLM Infrastructure & Observability Specialist โ helping teams monitor, debug, and control the cost of their AI systems in production. As AI features move from prototypes into production, teams quickly lose visibility into what their LLMs are actually doing: token spend, latency spikes, error rates, rate-limit throttling, and model drift often go unnoticed until they show up as a surprise invoice or an angry user. I bring production-grade observability engineering (Grafana, Prometheus/Mimir, Loki, Tempo) to AI workloads specifically. My background is deep infrastructure monitoring for cloud-native systems, which I now apply to: - LLM cost & token-usage dashboards (OpenAI, Azure OpenAI, Anthropic APIs) - Latency and error-rate monitoring for LLM/AI API calls, including gateway-level observability (Azure APIM as an LLM proxy) - Kubernetes-based AI/ML workload monitoring (GPU utilization, inference pod health, autoscaling behavior) - OpenTelemetry instrumentation for AI applications (traces across RAG pipelines, agent calls, vector DB queries) - Alerting for AI-specific failure modes: quota exhaustion, timeout spikes, degraded model responses - Infrastructure as Code (Terraform/OpenTofu) for reproducible AI observability stacks - CI/CD (GitLab, GitHub Actions) for deploying and versioning dashboards/alerts alongside AI infrastructure What sets me apart: I'm an infrastructure/SRE specialist who treats AI systems the same way I treat any other production system: with proper metrics, tracing, cost accountability, and on-call-ready alerting. Let's talk about what's actually happening inside your AI stack.