Back to LeaderboardLast snapshot: Oct 5
Mohit  P.
Badge

Mohit P.

Product Engineer at Big Tech | Re-Doc Founder | Technical Writer

Have been in this AI space for about 5 years now, even before the ChatGPT era :)) Have recently joined Upwork for side hustle. What I like: I have a passion for working specifically on building well curated datasets that can actually help llm to reason, think, question and reject things both in the Pre-training as well as Fine tuning steps. Work experience: I have worked extensively on building Data Extraction pipelines, creating fine tuning datasets for custom model training building and with integrating them with LLM models These pipelines are mostly on unstructured data sources like Scanned Documents, PDFs, Images, Email files etc. Dealing with unstructured data like PDFs, emails, Images etc., extracting information, maintaining the overall structure of the document and then passing to LLMs and retrieving results are one of the time intensive works. I specialize in this area, so can definitely help you out here. Projects: The Projects that I have worked on my company in recent 2 years (after the ChatGPT era) are based on OCR+LLM, PDF with images and table ingestion for RAG (Retrieval Augmented Generation) for longer context memorization, Agentic AI, Fine tuning LLM models on Custom Data, Creating Human in Loop pipelines for Feedback and post training (RLHF) and a little of Graph DB with Neo4j for replacing traditional RAG. I have listed some of the projects from my current CV -: โ€ข Gen AI based Email Helpdesk with automatic Database querying and validation โ€“ Worked on a Gen AI-based Email help desk responder where we automated steps like reading the incoming emails with questions from suppliers, understanding the supplier's ask, creating a SQL query based on the gathered information in the main database, and then finally responding to the email back with an appropriate answer based on domain knowledge and query outputs through Lang chain Agents/OpenAI function calling and custom prompting with multiple layers of API calls. โ€ข OCR + LLM based invoice entity extraction โ€“ Contributed to the development of a GenAI-based invoice entity extractor with table extraction that utilized cloud-based [GCP Document AI, Azure Form Recognizer, and AWS Textract] state-of-the art OCRs and combined them with the power of LLMs to extract information from invoices in various formats and templates. โ€ข Graph RAG for Automating Insurance Underwriting Guidelines Evaluation process โ€“ We developed a GenAI-based knowledge graph RAG (retrieval augmented generation) system where we utilized knowledge graphs to create nodes, triplets, and relationships for multiple PDFs filled with scanned images, text, and tables. Using this, we first extracted relevant information and then assessed it according to the underwriting guidelines of the insurance company. This enhancement led to an increase in the volume of submissions processed daily/monthly, while also reducing the time required to supply justifications and references for final decisions on each submission (explaining why the submission got approved or rejected). โ€ข Building Human in loop pipeline for Entity extraction workflow โ€“ Currently involved in building GenAI specific production required HIL pipelines with reusable components which can be leveraged for various GenAI Production Projects. Currently building for Invoices and Insurance domain specifically for the Entity extraction and Translation task. โ€ข Fine-Tuning LLMs on Domain Specific JSONL files โ€“ Currently involved in leveraging techniques like LORA, PEFT and as well as RLHF (Reinforcement Learning Human Feedback) for Finetuning Open (Llama, Qwen2.5, mistral etc.) as well Commercial models like GPT4o, Claude Sonnet, Gemini etc. on Insurance based data for task like RAG (Retrieval Augmented Generation), Underwriting etc. โ€ข Building Pipelines for Data Anonymization โ€“ Leveraging OCR models (open as well as Azure Doc Intelligence, GCP Document AI and Amazon Textract models) for Images / Scanned PDFs and Layout retainer PDF parsers along with Open Models like Qwen2.5 / Llama 3 for building Data Anonymization pipelines that can help prevent data leakages and help companies to be compliant. Please have a look at my LinkedIn Profile if you want to verify my identity, I can provide. (cannot add it on my profile bio, Upwork restricts)

๐Ÿ‡ฎ๐Ÿ‡ณHaldwani, India$60/hr100% JSS5.00 (3)Since Sep 2023
MRR
$85
๐ŸŒ#140K/398K
๐Ÿ‡ฎ๐Ÿ‡ณ#14.7K/43.2K
Recent Earnings
$511
๐ŸŒ#140K/398K
๐Ÿ‡ฎ๐Ÿ‡ณ#14.7K/43.2K
Total Earnings
$1,881
๐ŸŒ#324K/398K
๐Ÿ‡ฎ๐Ÿ‡ณ#35.9K/43.2K
Avg. Per Project
$376
๐ŸŒ#200K/398K
๐Ÿ‡ฎ๐Ÿ‡ณ#24.4K/43.2K
Recent Projects
1
2f0h
๐ŸŒ#128K/398K
๐Ÿ‡ฎ๐Ÿ‡ณ#14.4K/43.2K
Total Projects
5
6f0h
๐ŸŒ#253K/398K
๐Ÿ‡ฎ๐Ÿ‡ณ#31.3K/43.2K
MRR Performance Over Time
$50k$25k$0
6 mo ago3 mo agoNow
Coming SoonGathering historical data
World Skill Rankings
of 2
#1
Mixtral 8x7BTop <0.01%
#3Llama 2
#99Hugging Face
#393Machine Learning Model
#408ChatGPT API Integration
#459PyTorch
#1,133Prompt Engineering
#1,538Retrieval Augmented Generation
#1,669LangChain
#2,137Generative AI
#2,958OpenAI API
#10,046Python
๐Ÿ‡ฎ๐Ÿ‡ณIndia Skill Rankings
of 2
#1
Mixtral 8x7BTop <0.01%
#2Llama 2
#22Hugging Face
#57PyTorch
#72Machine Learning Model
#93ChatGPT API Integration
#165Prompt Engineering
#312Retrieval Augmented Generation
#353LangChain
#490Generative AI
#654OpenAI API
#1,732Python
Mohit P. โ€” Top 34% in India | UpworkMRR