Luca H.
LLM/VLM Engineer | Distributed Training, vLLM & CUDA Optimization
I help teams train, optimize, and deploy LLMs and multimodal models at production scale. I'm a Staff Machine Learning Engineer specializing in distributed LLM/VLM training, RL post-training, and GPU-optimized inference. I work across the full ML stack—from data pipelines and model architecture to multi-node training, evaluation, vLLM/TensorRT deployment, and performance profiling. Selected results: • Scaled distributed VLM training to 16 nodes / 128 GPUs through NCCL and eRDMA optimization • Increased average GPU utilization from approximately 30% to 85% and improved end-to-end vLLM inference throughput by 4.5x • Reduced multimodal video preprocessing time by more than 99% • Built a 32-GPU GRPO post-training pipeline that improved recall on key long-tail categories by up to 38% • Delivered a C++/CUDA TensorRT inference stack achieving 5 ms latency on an NVIDIA RTX 4090 • Contributed merged upstream improvements to LLaMA-Factory and LinkedIn's Liger Kernel I can help with: • LLM/VLM fine-tuning, SFT, LoRA/QLoRA, GRPO, and evaluation • Multi-node distributed training and GPU utilization optimization • vLLM, SGLang, and TensorRT inference pipelines • CUDA/Triton kernel integration, profiling, quantization, and latency optimization • Multimodal, video, computer-vision, and 3D AI systems Core stack: PyTorch, Hugging Face Transformers, LLaMA-Factory, DeepSpeed, verl, vLLM, TensorRT, CUDA, Triton, C++, and Python.