I'm a final-year B.Tech student specialising in AI/ML at IILM University, Greater Noida. I build things at the intersection of retrieval, LLMs, and computer vision, mostly RAG pipelines, document intelligence, and real-time detection systems.
Recent internships:
- AI/ML Intern @ Carnot Research Pvt. Ltd. (IITD-incubated) — built a production RAG + ETL platform for NITI Aayog's NDAP portal, integrated a 9,000+ line email/scheduling agent, and benchmarked 3 on-premise RAG architectures on Intel hardware
- Zee5 Tech & Innovation Centre — AI/ML work including YOLOv8 video pipelines and OmniDoc
- WESEE (Indian Navy defence R&D directorate) — LLM and RAG pipeline validation
Currently exploring distributed ML training, semi-supervised learning, and hardware-aware ML.
| Project | Description |
|---|---|
| Multiva.AI | Multilingual video dubbing platform that translates and re-voices video while preserving the speaker's own voice — Whisper transcription, NLLB-200 translation, XTTS v2 voice cloning; built with team ALCHEMISTS, 2nd place at a GGSIPU hackathon |
| VLM Business Card Lead Extraction | Turns photos of business cards into structured CRM-ready leads using a self-hosted Qwen3-VL vision-language model with grammar-constrained JSON output; FastAPI + PostgreSQL backend, React/TypeScript frontend, deployed on Azure — 96.8% field accuracy on real cards |
| Monocular RGB Sparse Point-Cloud SLAM | Reconstructs 3D camera trajectories and sparse point clouds from plain video using classical multi-view geometry (Lucas-Kanade optical flow, PnP-RANSAC, bundle adjustment) — no neural nets or GPU required, runs faster than real time on CPU |
