Skip to content
View irishz12's full-sized avatar

Block or report irishz12

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
irishz12/README.md

Hi, I'm Rishikesh

GenAI / Applied AI Engineer building retrieval, evaluation, and agentic systems in Python on AWS — focused on making AI systems verifiably correct, not just impressive in a demo.

Bengaluru, India · B.Tech CSE (AI & ML), JAIN University — 2026


Featured Projects

A cost-aware RAG system benchmarked on MultiHop-RAG (Tang & Yang, COLM 2024). An agentic controller iteratively retrieves across up to 3 hops until evidence is judged sufficient, evaluated against 4 baselines — Dense, Hybrid, Hybrid+Reranker, and a learned cost-optimized Adaptive router — using an independent LLM judge and a sealed, hash-verified holdout split. Scored: Evidence coverage 63.2% → 81.5% (baseline → agentic) on multi-hop questions · Adaptive RAG retains 80% answer quality at 21.8% lower cost 🔗 Live demo Python Qdrant FastAPI Next.js scikit-learn

Statistical release-gating for production LLM prompts. Scores a candidate prompt against production on quality, safety (real AWS Bedrock Guardrails), cost, and latency, runs paired bootstrap significance tests with Holm-Bonferroni correction, and returns a declarative GO / REVIEW / HOLD / INVALID verdict — no manual override. Scored: 5 of 6 real candidate prompts rejected or flagged for review before one passed both a 120-case dev set and a sealed 40-case holdout 🔗 Live demo Python FastAPI Next.js AWS Bedrock Guardrails

Generates synthetic tabular data (Gaussian Copula, CTGAN) and issues an ACCEPT / REGENERATE / REJECT decision across four evaluation pillars — fidelity, rare-case retention, privacy-risk, and downstream ML utility — verified against a cryptographically locked, SHA-256-hashed holdout split. 276 automated tests, 96% coverage. Scored: Gaussian Copula → REGENERATE (recall ratio 0.045 vs. 0.85 required) · CTGAN → ACCEPT (composite score 0.935), on the UCI Adult dataset (48,790 rows) Python pandas scikit-learn SDV Pandera

A classic supervised-learning predictive-maintenance pipeline — Random Forest vs. XGBoost vs. Logistic Regression under class-imbalance-aware training (3.4% failure rate), an F2-tuned decision threshold, 5-fold stratified cross-validation, and live per-prediction SHAP explanations served through FastAPI + Streamlit. Scored: PR-AUC 0.933 · 90.2% recall at a 0.6% false-positive rate on a held-out test set Python scikit-learn XGBoost SHAP FastAPI Docker

A full-stack RAG system answering factual and comparative questions across 6 insurance products straight from PDF brochures — hybrid BGE-M3 dense + BM25 sparse retrieval (RRF), BGE cross-encoder reranking, and Qwen3 32B generation via AWS Bedrock, with a citation-verification step that checks every generated claim against retrieved source text. Scored: Faithfulness 0.936 · Context Recall@4 0.837 · Answer Correctness 0.725 (RAGAS, 30-question benchmark) Python FastAPI Next.js Chroma AWS Bedrock

A LangGraph 4-agent pipeline (Planner → Investigation → Reflection Gate → Resolution) that proposes refund/replacement/escalation decisions — every proposal is validated by deterministic Python policy checks before execution. "LLM proposes, Python validates, Action executes." Scored: 100% issue-classification · 93.94% resolution accuracy · 96.97% final-status accuracy (33-scenario suite) Python LangGraph FastAPI Next.js Docker

An interactive demo exposing training-serving feature skew: a LightGBM fare model scored through duplicated ("Broken") vs. shared ("Correct") feature-transformation pipelines, across 9 features and 3 injectable skew scenarios. Showed: a distance-unit skew silently shifting a predicted fare from ~$20.29 (9/9 parity) to ~$29.19 (8/9 matched) Python LightGBM FastAPI React pytest

A pipeline processing 7.08M U.S. domestic flights (BTS, full-year 2024) through PostgreSQL bulk-loading and tail-number-based aircraft-rotation reconstruction, independently estimating delay propagation and benchmarking it against the BTS's official attribution. Found: 95.28% valid aircraft-link rate · 0.764 buffer-adjusted correlation · downstream delay rate falls from 69.06% to 15.29% as buffer time increases Python PostgreSQL SQL Plotly Dash


Other Projects


Tech Stack

GenAI & LLM: RAG · Multi-hop / agentic RAG · LangGraph · AWS Bedrock (Converse API, Bedrock Mantle, Guardrails) · Strands Agents · Local LLM serving (Ollama) · LLM evaluation (RAGAS, DeepEval, LLM-as-judge) · Statistical release gating (bootstrap CIs, Holm-Bonferroni) Cloud (AWS): Bedrock · EC2 · S3 · SageMaker AI · IAM · DynamoDB ML & MLOps: LightGBM · XGBoost · scikit-learn · SHAP · feature engineering · MLflow · Docker · GitHub Actions (CI/CD) Backend & Data: Python · SQL · FastAPI · Next.js · Pydantic · PostgreSQL · SQLAlchemy · pandas · NumPy · FAISS · Qdrant Visualization: Plotly Dash · Tableau · Power BI


Reach Me

LinkedIn Email

Pinned Loading

  1. Irishe1212z/GenAI-ReleaseGate Irishe1212z/GenAI-ReleaseGate Public

    Python 1

  2. multihop-rag-agent multihop-rag-agent Public

    Cost-aware agentic RAG with iterative multi-hop retrieval, benchmarked against 4 baselines using an independent LLM judge on MultiHop-RAG.

    Python 1

  3. Synthetic-Data-Validation-Optimization Synthetic-Data-Validation-Optimization Public

    Synthetic tabular data validation platform — deterministic accept/reject decisions backed by fidelity, privacy-risk, and ML-utility evaluation, with a cryptographically locked holdout.

    Python 1