Skip to content
View muditagrawal-alt's full-sized avatar
⚙️
Building
⚙️
Building

Block or report muditagrawal-alt

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
muditagrawal-alt/README.md
Mudit Agrawal

About Me

I'm a final-year B.Tech student specialising in AI/ML at IILM University, Greater Noida. I build things at the intersection of retrieval, LLMs, and computer vision, mostly RAG pipelines, document intelligence, and real-time detection systems.

Recent internships:

  • AI/ML Intern @ Carnot Research Pvt. Ltd. (IITD-incubated) — built a production RAG + ETL platform for NITI Aayog's NDAP portal, integrated a 9,000+ line email/scheduling agent, and benchmarked 3 on-premise RAG architectures on Intel hardware
  • Zee5 Tech & Innovation Centre — AI/ML work including YOLOv8 video pipelines and OmniDoc
  • WESEE (Indian Navy defence R&D directorate) — LLM and RAG pipeline validation

Currently exploring distributed ML training, semi-supervised learning, and hardware-aware ML.

Tech Stack

Python PyTorch OpenVINO HuggingFace OpenCV C Jupyter Git React

Featured Projects

Project Description
Multiva.AI Multilingual video dubbing platform that translates and re-voices video while preserving the speaker's own voice — Whisper transcription, NLLB-200 translation, XTTS v2 voice cloning; built with team ALCHEMISTS, 2nd place at a GGSIPU hackathon
VLM Business Card Lead Extraction Turns photos of business cards into structured CRM-ready leads using a self-hosted Qwen3-VL vision-language model with grammar-constrained JSON output; FastAPI + PostgreSQL backend, React/TypeScript frontend, deployed on Azure — 96.8% field accuracy on real cards
Monocular RGB Sparse Point-Cloud SLAM Reconstructs 3D camera trajectories and sparse point clouds from plain video using classical multi-view geometry (Lucas-Kanade optical flow, PnP-RANSAC, bundle adjustment) — no neural nets or GPU required, runs faster than real time on CPU

GitHub Stats

profile views

Pinned Loading

  1. OmniDoc OmniDoc Public

    An all purpose document llm made using RAG.

    Python 1

  2. Sliver-Smart-Video-Clipping-Tool Sliver-Smart-Video-Clipping-Tool Public

    Python 1

  3. Project-S.W.O.R.D Project-S.W.O.R.D Public

    Surveillance for Weapon Observation using Real-Time Deep Learning

    Jupyter Notebook

  4. Deepfake-and-Fake-News-Detector Deepfake-and-Fake-News-Detector Public

    Checks whether a news link, image or short video is what it claims to be. Local forensics plus an LLM fact-check, with the claims and sources shown.

    Python 3