Andrew G.
Senior AI/ML/Data Tech Lead | Strategy and Implementation
Technical Product Lead | LLM Systems & ML Infrastructure | MLOps AI product tech lead with a background in applied machine learning and systems architecture. Specializes in the productization of large language models (LLMs), recommendation systems, and computer vision pipelines. Fluent in the trade-offs between model fidelity, inference latency, compute cost, and statistical accuracy. Experienced in translating experimental research papers and half-trained checkpoints into production-grade, distributed systems. Operates at the intersection of prompt engineering, data ontology, and backend infrastructure. ###Technical Strategy & Systems Thinking - LLM Architecture Strategy: Roadmapping across fine-tuned OSS models (Llama 3, Mistral), closed-weight APIs (OpenAI, Anthropic), and hybrid routing layers. Defining context window utilization, retrieval-augmented generation (RAG) chunking strategies, and embedding model selection (e.g., Ada vs. Cohere vs. SBERT). - ML Evaluation & Validation: Designing offline/online evaluation frameworks beyond accuracy—specializing in hallucination rate, perplexity, toxicity filters, and adversarial robustness. Experience with human-in-the-loop (HITL) labeling workflows and active learning loops. - Infrastructure & MLOps: Defining requirements for feature stores, model registries, and inference orchestration. Deep understanding of GPU/TPU utilization, autoscaling policies, and cold-start mitigation. Familiar with Kubernetes, Ray, and vector database sharding strategies. - Data-Centric AI: Prioritizing data curation over architecture tweaks. Expertise in synthetic data generation, class imbalance correction, and weak supervision (Snorkel/Skweak) for low-resource domains. ###Technical Stack & Implementation Fluency - Languages & Querying: Python (scripting, data analysis), SQL (complex aggregations, feature engineering), GraphQL/REST. - Frameworks & Libraries: LangChain, LlamaIndex, Hugging Face Transformers, PyTorch, TensorFlow, Scikit-learn, spaCy. - Infrastructure & Tooling: AWS SageMaker, Bedrock; GCP Vertex AI; Databricks; Weights & Biases; MLflow; Docker; Kubernetes. - Vector/NoSQL: Pinecone, Milvus, Chroma, Redis, PostgreSQL (pgvector). - Experimentation: A/B testing with inference shadows, canary deployments, multi-armed bandit algorithms. ###Technical Implementation Highlights 1. RAG Architecture for Enterprise Q&A *Led product requirements for a retrieval-augmented generation system targeting legal document analysis. Evaluated trade-offs between dense (DPR, ColBERT) and sparse (BM25) retrievers. Defined chunking ontology and metadata filtering schemas to reduce latency from 2.3s to 480ms while maintaining top-3 recall above 89%. Implemented re-ranking layer to improve answer relevancy by 32%.* 2. LLM Fine-Tuning & Cost Optimization *Managed roadmap for domain-adaptation fine-tuning of a 7B parameter model. Instrumented LoRA vs. full fine-tuning experiments to optimize for VRAM constraints. Reduced inference cost per 1M tokens by 63% through quantization (bitsandbytes) and speculative decoding implementation. Collaborated with MLEs to build a feedback loop using human preference data for DPO training.* 3. Real-Time Computer Vision Pipeline *Shipped an on-device object detection model for mobile IoT. Defined requirements for model compression (TensorFlow Lite, pruning) to meet <15MB size constraint and 30 FPS threshold on edge devices. Implemented frame-sampling strategy to reduce cloud egress costs by 73%.* 4. Model Observability & Drift Detection Implemented statistical alerting for model performance degradation in production. Defined thresholds using PSI (Population Stability Index) and KL divergence on embedding distributions. Built requirements for explainability layer using SHAP and LIME to debug failure modes flagged by trust & safety teams.