A live, interactive version of the US Census Income Predictive Model originally built in Orange Data Mining. Same data, same preprocessing, same Gradient Boosting classifier, rebuilt in Python so it can run as a deployed web app instead of static report screenshots.
Try it live: [add your Streamlit Cloud URL here after deploying]
- Predict: enter age, hours worked, education, occupation, and demographics, get a live income bracket prediction with a SHAP breakdown of what drove it
- Explore the Data: interactive versions of the report's income-by-education, income-by-sex, and age-vs-income visuals
- Income, Education & Voting: the state-level 2020 election join from the report, income/education vs Biden vote share by coastal/inland state
| Report step | This app |
|---|---|
| Data Sampler (5,000 rows) | preprocess.py, same random sample size |
| Edit Domain / Formula widgets (recoding sex, race, PoB, CoW, marital, education, occupation) | preprocess.py, same category buckets |
| Outlier removal (income > $500K, hours > 80/wk) | preprocess.py |
| Test & Score (7 classifiers, Gradient Boosting best: AUC 0.879, acc 0.787) | train_model.py, Gradient Boosting only (test acc ~0.78 on this split) |
| SHAP explanation | app.py, live per-prediction SHAP plot instead of a static beeswarm |
| Geo map + t-tests (income/education vs 2020 vote) | build_voting.py + app.py, interactive scatter plots instead of static Geo Map widgets |
Full statistical detail (Welch's t-tests, Bayesian model comparison, Zipf/Pareto analysis) is in the original report, linked above — this app focuses on the parts that benefit from being interactive.
pip install -r requirements.txt
streamlit run app.pyThe data/ folder already contains a processed sample, trained model, and state voting table. To regenerate them from raw census data:
python preprocess.py # samples + cleans Census_Data.csv
python build_voting.py # joins with 2020 election results
python train_model.py # trains Gradient Boosting + saves SHAP explainerYou'll need Census_Data.csv, Attribute_Values.csv, and voting_2020.csv (not included in this repo due to size — see the original report repo).
- Push this folder to a new GitHub repo
- Go to share.streamlit.io, sign in with GitHub
- Point it at this repo,
app.pyas the entry file - Free tier, deploys in a couple of minutes, gives you a public
.streamlit.appURL
Python, scikit-learn (Gradient Boosting), SHAP, Streamlit, Plotly