Musharaf A.
Generative AI Engineer | Production RAG, LLM Agents & Vector Search
I build RAG systems and LLM agents that hold up in production. On my last enterprise conversational AI platform, retrieval accuracy went up 37% and response latency dropped by 17 seconds after I rebuilt the chunking and retrieval layer. I'm Lead AI Engineer at a London-based company and take a small number of Upwork contracts alongside it; most recently an 18-month, 716-hour engagement that one client kept renewing. WHAT I BUILD - RAG and conversational AI over private documents: hierarchical, recursive and topic-based chunking, hybrid retrieval (BM25 + dense vectors), re-ranking, sensitive-data guards, and multi-provider LLM failover - LLM agents with real tools: SQL, graph queries, web search, internal APIs using LangGraph, MCP and n8n - Vector database architecture โ Milvus, Pinecone, Chroma, MongoDB Atlas, Elasticsearch, index design, quantization, query cost tuning - LLMOps: embedding pipelines, CI/CD for models, MLflow, monitoring, drift detection, inference autoscaling on Azure, GCP and AWS - No-code ML platforms so non-technical teams can upload data, tune models and ship predictions without engineering support RESULTS FROM RECENT WORK Enterprise Conversational AI Platform: Rebuilt chunking and retrieval for a document-grounded assistant, added sensitive-data guards and multi-provider LLM switching. Retrieval accuracy +37%, response latency โ17s. LangChain, LangGraph, MongoDB Atlas, Azure AI Foundry. Crypto Transaction Intelligence: Modeled entities in Memgraph with sanctions screening and IP/VPN detection, and enabled LLM-to-Cypher natural language querying. Cut investigation time from days to hours. Memgraph, LLMs, Binance/OKX/Kraken APIs. Logistics ML: Packing-box selection model that improved packing efficiency 32% and reduced material waste; distributed Random Forest ETA predictor that improved delivery-time accuracy 18%. PySpark, scikit-learn, distributed training. Medical Text Summarization API: Fine-tuned Pegasus for long-form clinical text with multi-GPU distributed training, deployed as a real-time inference API. Pegasus, Hugging Face, multi-GPU. BACKGROUND MS in Computer Science (Information Technology University, Pakistan): Computer Vision, Deep Learning, Big Data. Gold Medalist in BS Information Technology. Fully funded PEEF graduate scholarship. Five years across FinTech and crypto, healthcare, logistics, e-commerce, media analytics and sports. HOW WE START Message me with what you are building and what's blocking you. I'll reply within a few hours with an honest read on whether it's a fit, including when it isn't, and who you should talk to instead.