Skip to content

Latest commit

 

History

17 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Census Income Model — Interactive App

A live, interactive version of the US Census Income Predictive Model originally built in Orange Data Mining. Same data, same preprocessing, same Gradient Boosting classifier, rebuilt in Python so it can run as a deployed web app instead of static report screenshots.

Try it live: [add your Streamlit Cloud URL here after deploying]

What it does

  • Predict: enter age, hours worked, education, occupation, and demographics, get a live income bracket prediction with a SHAP breakdown of what drove it
  • Explore the Data: interactive versions of the report's income-by-education, income-by-sex, and age-vs-income visuals
  • Income, Education & Voting: the state-level 2020 election join from the report, income/education vs Biden vote share by coastal/inland state

How it maps to the original report

Report step This app
Data Sampler (5,000 rows) preprocess.py, same random sample size
Edit Domain / Formula widgets (recoding sex, race, PoB, CoW, marital, education, occupation) preprocess.py, same category buckets
Outlier removal (income > $500K, hours > 80/wk) preprocess.py
Test & Score (7 classifiers, Gradient Boosting best: AUC 0.879, acc 0.787) train_model.py, Gradient Boosting only (test acc ~0.78 on this split)
SHAP explanation app.py, live per-prediction SHAP plot instead of a static beeswarm
Geo map + t-tests (income/education vs 2020 vote) build_voting.py + app.py, interactive scatter plots instead of static Geo Map widgets

Full statistical detail (Welch's t-tests, Bayesian model comparison, Zipf/Pareto analysis) is in the original report, linked above — this app focuses on the parts that benefit from being interactive.

Running locally

pip install -r requirements.txt
streamlit run app.py

Rebuilding the data/model from scratch

The data/ folder already contains a processed sample, trained model, and state voting table. To regenerate them from raw census data:

python preprocess.py    # samples + cleans Census_Data.csv
python build_voting.py  # joins with 2020 election results
python train_model.py   # trains Gradient Boosting + saves SHAP explainer

You'll need Census_Data.csv, Attribute_Values.csv, and voting_2020.csv (not included in this repo due to size — see the original report repo).

Deploying

  1. Push this folder to a new GitHub repo
  2. Go to share.streamlit.io, sign in with GitHub
  3. Point it at this repo, app.py as the entry file
  4. Free tier, deploys in a couple of minutes, gives you a public .streamlit.app URL

Stack

Python, scikit-learn (Gradient Boosting), SHAP, Streamlit, Plotly

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages