An end-to-end machine learning web application that predicts house prices in real time. Built from a raw Kaggle dataset all the way to a deployed, production-style web app — including data cleaning, model comparison, a FastAPI backend, and a React + TypeScript frontend.
This project estimates residential property prices in India based on features like location, carpet area, number of bathrooms, furnishing status, and more. Four regression models were trained and compared on the same cleaned dataset, and the best-performing lightweight model (XGBoost) powers the live estimator, alongside two additional models the user can choose from for comparison.
| Layer | Technology |
|---|---|
| Data cleaning & modeling | Python, Pandas, NumPy, scikit-learn, XGBoost |
| Backend API | FastAPI, Pydantic, Uvicorn |
| Frontend | React, TypeScript, Vite |
| Charts | Recharts |
| Icons | react-icons |
Source: House Price by Juhi Bhojani on Kaggle (~187,000 real property listings from India).
The raw CSV is not committed to this repository (see .gitignore). To reproduce the notebook:
pip install kaggle
# Get your API token from Kaggle → Settings → API → "Create New Token"
# Place kaggle.json in ~/.kaggle/ (or C:\Users\<you>\.kaggle\ on Windows)
kaggle datasets download -d juhibhojani/house-price -p notebooks/data --unzipFour models were trained and evaluated on the same held-out test set (20% split):
| Model | MAE | RMSE | R² |
|---|---|---|---|
| Random Forest | ₹10.4L | ₹42.8L | 0.909 |
| XGBoost | ₹16.0L | ₹48.5L | 0.883 |
| Gradient Boosting | ₹24.1L | ₹54.8L | 0.851 |
| Linear Regression | ₹44.6L | ₹334.8L | −4.567 |
Random Forest achieved the highest accuracy, but its serialized model size (~467 MB) made it impractical to ship in this repository and deploy easily. XGBoost was chosen as the production model — it keeps 97% of Random Forest's accuracy at a fraction of the size (~1.2 MB), making it the better engineering trade-off for a deployed app. Linear Regression's negative R² shows the pricing pattern in this data is too non-linear for a simple linear model — it's kept in the comparison for that reason, and users can still try it live in the estimator.
Full training, cleaning steps, and evaluation plots are in notebooks/data/house_price_model.ipynb.
cd backend
python -m venv .venv
.venv\Scripts\activate # Windows
# source .venv/bin/activate # macOS/Linux
pip install -r requirements.txt
cp .env.example .env
uvicorn app.main:app --reloadThe API will be available at http://localhost:8000, with interactive docs at http://localhost:8000/docs.
| Variable | Description | Example |
|---|---|---|
MODEL_PATH |
Path to the default model file | models/house_price.pkl |
LOCATIONS_PATH |
Path to allowed locations list | models/locations.json |
ALLOWED_ORIGIN |
CORS-allowed frontend origin | http://localhost:5173 |
GET /health — health check
curl http://localhost:8000/healthGET /models — list available models
curl http://localhost:8000/modelsPOST /predict — get a price estimate
curl -X POST http://localhost:8000/predict \
-H "Content-Type: application/json" \
-d '{
"location": "other",
"carpet_area_sqft": 1200,
"floor_num": 3,
"bathroom": 2,
"balcony": 1,
"furnishing": "Furnished",
"transaction": "Resale",
"ownership": "Freehold",
"facing": "East",
"model": "xgboost"
}'Response:
{
"predicted_price": 7499999.99,
"model_used": "xgboost"
}cd backend
python -m pytestcd frontend
npm install
cp .env.example .env
npm run devThe app will be available at http://localhost:5173.
| Variable | Description | Example |
|---|---|---|
VITE_API_BASE_URL |
Base URL of the backend API | http://localhost:8000 |
npm run build(Add 2–3 screenshots here of the Home, Estimator, and Result pages)




