Gregory H.
Ex-Meta,Twitter | Senior AI Engineer | Full Stack | LLMs, RAG, Claude
I'm an AI architect who spent a decade building world-class engineering teams at Facebook, Twitter, and Aurora—then pivoted to building the software itself. Now I design and ship full-stack, AI-powered products end-to-end: from GPU-accelerated ML pipelines to polished user experiences. I believe the best AI feels invisible while its impact remains unmistakable. SPECIALIZED AI EXPERTISE RAG & Knowledge Systems • Built production RAG chatbots using pgvector embeddings, structured metadata, and conversation context • Designed knowledge graphs (Characters, Locations, Factions, Items) with alias normalization and semantic entity linking • Vector DBs: Pinecone, pgvector, Chroma | Frameworks: LangChain, LlamaIndex Speech-to-Text & Audio AI • GPU-accelerated STT pipelines using CUDA-enabled Whisper Large-v3-Turbo on RunPod clusters • Multi-stage LLM correction: NER correction → speaker attribution → narrative stitching • Real-time streaming transcription with chunk-level hashing and retry semantics Agentic AI & LLM Systems • Multi-stage LLM pipelines for entity extraction, summarization, and content generation • LLM-judge validation layers for dataset QA and fine-tuning data integrity • Prompt engineering and hallucination mitigation with source attribution AI Image Generation • Stable Diffusion deployment on GPU workers with queue-aware batch scheduling • Custom prompt engineering for character portraits and concept art generation TECHNICAL STACK Generative AI & LLMs • Models: GPT-5, Claude 4.5, Whisper Large-v3, Faster-Whisper, Stable Diffusion • Fine-tuning: LoRA, instruction tuning, LLM-judge validation • Serving: RunPod GPU clusters, custom autoscalers, queue-aware batch scheduling ML & Deep Learning • Frameworks: PyTorch, sentence-transformers, Hugging Face Transformers • NLP: Named entity recognition, speaker diarization, semantic entity extraction • Embeddings: pgvector, vector search, RAG pipelines Programming & Development • Languages: TypeScript/JavaScript, Python, SQL • Backend: Node.js, FastAPI, Prisma ORM • Frontend: React 18, Next.js 14 (App Router), MUI, Tailwind CSS Cloud & Infrastructure • Compute: Vercel, DigitalOcean Kubernetes, RunPod GPU, Railway • Storage: PostgreSQL/Neon, Cloudflare R2, Redis Streams • DevOps: Docker, Kubernetes, Prometheus/Grafana, GitHub Actions Data Engineering • Streaming: Redis Streams, queue-based job orchestration • Databases: PostgreSQL, pgvector, Neo4j • Pipelines: Real-time audio ingestion, distributed ASR job scheduling CORE COMPETENCIES Distributed Systems • Microservices orchestration across Vercel, Kubernetes, Railway, and GPU pods • Redis Stream + Ledger orchestration for ingestion, backpressure, idempotency • Fault-tolerant autoscaling based on queue depth and GPU utilization Production AI • GPU cluster management with dynamic provisioning/tear down • Model serving optimization and real-time inference pipelines • 99%+ reliability with graceful degradation Full-Stack Ownership • Database schema design with RAG-ready embeddings • API development and polished user interfaces • Subscription management and real-time data syncing KEY PROJECT: ARCHIVIST AI Co-created a large-scale AI platform for TTRPG players integrating real-time STT, RAG, AI image generation, and structured worldbuilding knowledge graphs. GPU-Accelerated Transcription Pipeline • CUDA-enabled Whisper on RunPod with custom autoscaling logic and concurrency management • Simul-Streaming gateway with chunk-level job hashing, retry semantics, and guardrails • Multi-hour Discord session ingestion via custom bot + Opus frame capture + R2 storage Intelligent Knowledge Graph • Prisma + Neon Postgres schema with worldbuilding entities and unique graph linking rules • Semantic entity extraction with alias heuristics and normalization • AI-powered session summaries, beat tracking, and smart compendium search Production SaaS Platform • Full Next.js App Router application with dashboards, campaign tools, and subscriptions • Microservices environment with Prometheus/Grafana observability • Real-time transcript assembly and data syncing Results • Eliminated manual note-taking for thousands of hours of gameplay • Single source of truth for complex, long-running campaigns • 99%+ transcription reliability with graceful error handling MY APPROACH I architect intelligent systems that learn, adapt, and scale: • Rapid Iteration: MVP in days, production in weeks • Systems Thinking: Designed for scale, reliability, and graceful degradation • Full Ownership: Infrastructure, ML, backend, frontend—from first line to production metrics • AI That Works: Production-grade pipelines, not just demos Whether you need RAG systems, speech-to-text pipelines, LLM integrations, or end-to-end AI products, I bring engineering rigor and product sensibility to every project. Ready to build something real? Let's talk.