An AI-powered wine intelligence platform that uses machine learning to analyze wine chemistry, predict quality, discover wine profiles, and explain predictions with Google Gemini.
VinoLens is a full-stack machine learning project built around the Wine Quality dataset. It combines classical machine learning with a modern web application and generative AI.
The goal isn't simply to predict a wine's quality score โ VinoLens aims to help users understand why a wine receives its prediction and discover patterns across different wine profiles.
Enter the physicochemical properties of a wine and receive a predicted quality score.
- Linear Regression / other ML models
- Model evaluation and comparison
- Quality prediction from wine chemistry
- Feature analysis
Explore the dataset interactively.
- Visualize relationships between wine properties
- Filter wines by characteristics
- Explore quality distributions
- Compare different wine profiles
Use unsupervised learning to discover naturally occurring groups of wines.
Using clustering algorithms such as K-Means, VinoLens can identify wine profiles based on their chemical characteristics without using the quality label.
Understand what influences predictions.
- Feature importance
- Feature correlations
- Prediction analysis
- Model performance metrics
Google Gemini provides a natural-language explanation of the ML model's prediction.
Instead of simply showing:
Predicted Quality: 7.2
VinoLens can explain:
The wine's relatively high alcohol and sulphate levels contribute positively to the predicted quality, while volatile acidity has a negative influence.
The ML model makes the prediction. Gemini explains it.
โโโโโโโโโโโโโโโโโโโ
โ Next.js โ
โ Frontend โ
โโโโโโโโโโฌโโโโโโโโโ
โ
HTTP / REST
โ
โผ
โโโโโโโโโโโโโโโโโโโ
โ FastAPI โ
โ Backend โ
โโโโโโโโโโฌโโโโโโโโโ
โ
โโโโโโโโโโโโโโโผโโโโโโโโโโโโโโ
โ โ โ
โผ โผ โผ
โโโโโโโโโโโโ โโโโโโโโโโโโ โโโโโโโโโโโโ
โ ML Model โ โ K-Means โ โ Gemini โ
โRegressionโ โClusteringโ โ AI โ
โโโโโโโโโโโโ โโโโโโโโโโโโ โโโโโโโโโโโโ
โ โ โ
โโโโโโโโโโโโโโโผโโโโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโ
โ Wine Quality โ
โ Dataset โ
โโโโโโโโโโโโโโโโโโโ
- Python
- Pandas
- NumPy
- Scikit-learn
- Matplotlib
- Jupyter Notebook
- Python
- FastAPI
- Pydantic
- Uvicorn
- Next.js
- React
- TypeScript
- Tailwind CSS
- Google Gemini API
- Git
- GitHub
- Docker
- Docker Compose
vinolens/
โ
โโโ frontend/ # Next.js 14+ (Deployed to Vercel)
โ โโโ app/ # App Router (pages & layouts)
โ โโโ components/ # UI & chart components
โ โ โโโ ui/ # Reusable primitives (buttons, inputs)
โ โ โโโ wine-form.tsx # Input chemical properties
โ โ โโโ prediction-card.tsx # Prediction & Gemini explanation display
โ โโโ lib/ # API client (fetch to FastAPI backend)
โ โโโ public/ # Static assets (images, icons)
โ โโโ package.json
โ โโโ tailwind.config.ts
โ โโโ tsconfig.json
โ
โโโ backend/ # FastAPI Application (Deployed to Render / Fly.io / Docker)
โ โโโ app/
โ โ โโโ main.py # App factory, CORS, middleware, router mount
โ โ โโโ config.py # Pydantic BaseSettings (GEMINI_API_KEY, MODEL_PATH)
โ โ โโโ api/
โ โ โ โโโ deps.py # Shared dependencies (model loader, auth)
โ โ โ โโโ v1/
โ โ โ โโโ router.py # API v1 aggregator
โ โ โ โโโ endpoints/
โ โ โ โ โโโ predict.py # POST /predict (Quality score & feature impact)
โ โ โ โ โโโ cluster.py # POST /cluster (Wine profile discovery)
โ โ โ โ โโโ wines.py # GET /wines (Dataset exploration & stats)
โ โ โ โ โโโ explain.py # POST /explain (Gemini sommelier reasoning)
โ โ โโโ schemas/ # Pydantic v2 schemas
โ โ โ โโโ wine.py # WineFeaturesInput, WineStats
โ โ โ โโโ prediction.py # PredictionResponse, ExplanationResponse
โ โ โโโ services/ # Core business logic
โ โ โ โโโ predictor.py # Loads sklearn pipeline and computes inference
โ โ โ โโโ clusterer.py # Unsupervised K-Means clustering service
โ โ โ โโโ gemini_sommelier.py # Google Gemini API prompt & explanation engine
โ โ โโโ models/ # Production model artifacts loaded by FastAPI
โ โ โโโ quality_regressor.joblib
โ โ โโโ wine_clusters.joblib
โ โ
โ โโโ Dockerfile # Lightweight FastAPI container for deployment
โ โโโ .dockerignore
โ
โโโ ml/ # Offline ML Training & Research
โ โโโ notebooks/ # Exploratory & Prototyping
โ โ โโโ 01_eda.ipynb # Feature distributions & correlations
โ โ โโโ 02_regression.ipynb # Quality prediction experiments
โ โ โโโ 03_clustering.ipynb # K-Means & PCA profiling
โ โโโ src/ # Reproducible pipeline code
โ โ โโโ __init__.py
โ โ โโโ data.py # Data loaders & validation
โ โ โโโ features.py # Scaling, transforms & feature engineering
โ โ โโโ train_regression.py # Model training script -> saves to backend/app/models/
โ โ โโโ train_clustering.py # Clustering script -> saves to backend/app/models/
โ โ โโโ evaluate.py # MSE, RMSE, R2, Silhouette scores
โ โโโ configs/
โ โโโ hyperparameters.yaml # Configurable training params
โ
โโโ data/ # Data storage (gitignored except raw sample)
โ โโโ raw/ # Immutable raw data (WineQT.csv)
โ โโโ processed/ # Cleaned / split / scaled datasets
โ
โโโ tests/ # Automated Test Suite
โ โโโ backend/
โ โ โโโ test_api.py # FastAPI TestClient endpoint tests
โ โ โโโ test_predictor.py # Inference & Gemini mock tests
โ โโโ ml/
โ โโโ test_data_pipeline.py # Data leakage & shape tests
โ
โโโ .github/ # CI/CD Workflows
โ โโโ workflows/
โ โโโ test.yml # Linting & unit tests on PR
โ โโโ deploy.yml # Automated deployment
โ
โโโ .env.example # Documented environment variables
โโโ .gitignore
โโโ docker-compose.yml # Orchestrates Frontend + Backend locally
โโโ pyproject.toml # Single unified dependency management with uv
โโโ README.md
VinoLens uses the Wine Quality dataset containing physicochemical measurements of wine along with a quality score.
The model can use features such as:
| Feature | Description |
|---|---|
fixed acidity |
Fixed acids in the wine |
volatile acidity |
Volatile acidity |
citric acid |
Citric acid concentration |
residual sugar |
Remaining sugar |
chlorides |
Salt concentration |
free sulfur dioxide |
Free sulfur dioxide |
total sulfur dioxide |
Total sulfur dioxide |
density |
Density of the wine |
pH |
Acidity level |
sulphates |
Sulphate concentration |
alcohol |
Alcohol percentage |
quality
The quality score is an integer rating assigned to each wine.
VinoLens explores the dataset through multiple machine learning approaches.
The initial model treats wine quality as a regression problem:
Wine Chemistry
โ
โผ
Machine Learning Model
โ
โผ
Predicted Quality
Example:
Input โ Wine chemical properties
Output โ 7.2 / 10
Future versions can also treat the problem as classification:
quality >= 7 โ Good
quality < 7 โ Not Good
VinoLens also explores the dataset without using the quality label.
K-Means clustering can identify groups of wines with similar chemical characteristics.
Wine Data
โ
โผ
K-Means
โ
โโโ Cluster 1
โโโ Cluster 2
โโโ Cluster 3
This allows VinoLens to answer questions such as:
"What types of wines naturally occur in this dataset?"
Google Gemini is used as an explanation layer, rather than replacing the machine learning model.
Wine Features
โ
โผ
ML Model
โ
โโโ Prediction
โโโ Feature influence
โโโ Cluster
โ
โผ
Gemini
โ
โผ
Human-readable explanation
This separation keeps the ML prediction deterministic and allows Gemini to focus on communicating the results.
- Obtain Wine Quality dataset
- Load dataset with Pandas
- Explore statistics
- Check missing values
- Visualize distributions
- Analyze feature correlations
- Prepare features and target
- Train/test split
- Implement linear regression
- Evaluate MSE
- Evaluate RMSE
- Evaluate Rยฒ
- Experiment with different features
- Try additional regression models
- Compare model performance
- Implement classification
- Implement K-Means clustering
- Investigate feature importance
- Improve preprocessing
- Create FastAPI application
- Create
/predictendpoint - Create
/clustersendpoint - Create
/statisticsendpoint - Load trained ML models
- Add request validation
- Build landing page
- Create wine analysis form
- Build prediction dashboard
- Add interactive charts
- Build wine explorer
- Add clustering visualization
- Add responsive design
- Integrate Gemini API
- Generate prediction explanations
- Explain influential features
- Generate wine profile summaries
- Dockerize application
- Add environment configuration
- Add automated tests
- Deploy frontend
- Deploy backend
- Document API
- Add CI/CD
1. User opens VinoLens
โ
2. Selects "Analyze Wine"
โ
3. Enters wine chemistry
โ
4. FastAPI receives request
โ
5. ML model generates prediction
โ
6. Model determines wine profile
โ
7. Gemini explains the result
โ
8. User sees:
Quality: 7.2 / 10
Profile: High-Alcohol Balanced
+ Positive factors
- Negative factors
๐ก AI explanation
Potential future improvements include:
- Wine recommendation system
- Similar-wine search
- "Find wines similar to this one"
- User accounts and saved analyses
- Prediction history
- Wine comparison
- Interactive what-if analysis
- SHAP-based model explanations
- Experiment tracking
- Model versioning
- Additional wine datasets
- Red vs. white wine analysis
- Personalized wine recommendations
VinoLens is an educational machine learning project.
The predictions represent patterns learned from the dataset and should not be interpreted as professional wine certification, laboratory analysis, or expert sommelier judgment.
This project is designed to demonstrate practical understanding of:
- Supervised learning
- Regression
- Classification
- Unsupervised learning
- Clustering
- Feature engineering
- Model evaluation
- Data visualization
- REST APIs
- Full-stack development
- Generative AI integration
- ML model deployment
๐ง Currently in development
The project is being developed incrementally alongside the study of machine learning concepts.
Data exploration โ Linear Regression โ Model Evaluation
This project is intended for educational and portfolio purposes.