I turn research into production β systems that see, listen, reason, and create.
AI Engineer with 2 years of experience building and shipping production ML systems across Computer Vision, OCR, LLM / Agentic AI, and Speech (ASR). I own the full lifecycle β fine-tuning (PEFT/LoRA, 4-bit), evaluation, and GPU-optimized inference β and enjoy taking a model from a research paper all the way to a reliable, end-to-end pipeline.
- π Currently building OCR, object-detection, and LLM-agent systems @ WorkerBot AI
- π§ Deepest expertise: speech recognition & efficient LLM fine-tuning
- β‘ Fun fact: I fine-tuned Gemma 3N for Vietnamese ASR down to 7.21% WER
End-to-end Vietnamese speech recognition on a fine-tuned Gemma 3N β built from scratch.
- π― 7.21% WER on a 5,000-sample test set (0 empty predictions, ~97K reference words)
- π§© Production inference pipeline:
Demucs β denoise β VAD β overlap-aware chunking β context-aware decoding- βοΈ PEFT/LoRA + 4-bit quantization via Unsloth β trainable on a single consumer GPU
- π¦ Clean, reproducible codebase: separate
train/evaluate/predictmodules
| Project | What it does | Tech |
|---|---|---|
| ποΈ Audio2Text | Vietnamese ASR toolkit on fine-tuned Gemma 3N β training, eval & production inference. 7.21% WER | Gemma 3N PEFT/LoRA Unsloth Demucs VAD |
| π StoryForge | Multi-agent story generator β 13-agent drama simulation, LLM-as-judge auto-revision & RAG | FastAPI LLM Multi-Agent RAG |
| π§βπΌ AI HR Interview | Full-stack AI interviewer with real-time voice/video via Gemini Live + JDβCV matching | Next.js Gemini Live PostgreSQL Redis |
| π MedGraph | Drug-interaction cascade analyzer β knowledge graph over CYP450 pathways on real FDA data | FastAPI React Knowledge Graph |
| π’ Date-Recognition | Expiry-date OCR β YOLOv8 detection + CRNN/CTC recognition with Streamlit UI | YOLOv8 OCR CRNN |
| π FaceReg | Real-time face recognition β MTCNN + FaceNet across image, video & live camera | PyTorch MTCNN FaceNet |
Languages
ML / LLM
Computer Vision Β· Speech
Backend Β· Infra
- Efficient LLM fine-tuning & on-device / low-VRAM inference
- Multi-agent systems and autonomous research workflows
- Advanced OCR & document understanding
"Turn research into systems people can actually use."
π« Reach me at nt.hieu2207@gmail.com Β· β Star anything you find useful!