Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

US Census Income Predictive Model

End-to-end data mining analysis of a 1.66M-row US Census dataset, built in Orange Data Mining. Covers preprocessing, exploratory analysis, hypothesis testing, and classification, with a focus on identifying which factors actually predict income.

Tools: Orange Data Mining, SHAP, Welch's t-test, Gradient Boosting, Bayesian model comparison


Overview

Starting from a 1.66 million row census file, this project builds a full pipeline: sampling, cleaning, feature engineering, exploratory analysis, statistical testing, and predictive modelling, then extends into a geo-spatial analysis linking income and education to 2020 US election results.

The workflow is built entirely in Orange's visual programming environment rather than a Python/pandas script. The full write-up, including all figures and reasoning, is in report.pdf. This README summarises the approach and findings.

Pipeline

1. Preprocessing

  • Randomly sampled 5,000 of 1.66M rows to reduce computation while preserving population distribution
  • Converted state from numeric to categorical via a lookup table
  • Recoded sex, race, place of birth, class of worker, and marital status into simplified categories using the Formula widget
  • Grouped education into no diploma / high school / post-high-school / degree bands, and occupation into SOC-based industry groups
  • Removed outliers (415 rows, 8.3%), primarily extreme income (>$500K) and hours (>80/week) values, reducing mean income skew from $56,108 to $54,785

2. Exploratory analysis

  • Confirmed strong right-skew in income (typical of income data), corrected with a log(income + 1) transform
  • Verified the skew is Pareto-like: a log(rank) vs log(income) plot showed slope ≈ -1.0, meaning the top 1% of earners hold a disproportionate share of total income
  • Ran Welch's t-tests across sex, race, and place of birth:
    • Sex: $22,743 gap (t = 11.72, p < 0.0001)
    • Race: $8,351 gap (t = 3.74, p < 0.0002)
    • Place of birth: no significant effect (t ≈ 1.0, p ≈ 0.3)

3. Predictive modelling

  • Split income at the median ($38,010) into low/high subgroups
  • Compared 7 classifiers; Gradient Boosting performed best (AUC 0.879, accuracy 0.787), confirmed via Bayesian pairwise comparison (0.875 probability of outperforming Random Forest)
  • Used SHAP values and permutation importance to explain predictions: hours worked, occupation, education, and age were the strongest predictors, education included

Key finding: although education correlates with income on its own (r = 0.67), it drops to a secondary factor once hours worked and occupation are accounted for. Time on the job and industry sector matter more than years of schooling.

4. Geo-spatial extension

  • Joined census income/education data with 2020 election results by state
  • States with higher income and education aligned closely with Biden-voting states
  • Education showed a significant link to voting outcome (t = 2.88, p < 0.01); income alone did not (t = -1.92, p ≈ 0.10)
  • Split states into coastal/inland: income predicted Democratic vote share more strongly in coastal states (r = 0.44) than inland (r = -0.01), though the difference fell just short of significance (Fisher z = 1.598, p = 0.055)

Limitations

  • 5,000-row sample may under-represent rare subgroups and extreme high earners present in the full 1.66M-row dataset
  • Binary groupings (white/non-white, male/female, coastal/inland) simplify analysis but remove within-group variation and intersectional effects
  • Aggregate state-level correlations do not imply individual-level causation
  • DC and Puerto Rico were excluded due to data alignment issues; DC's exclusion likely weakens the income/education-voting correlations, since it is a high-income, high-education, heavily Democratic outlier
  • "Coastal" as a binary category is a rough proxy that likely captures urbanicity more than geography itself

Files

  • report.pdf: full write-up with all figures, workflows, and statistical detail
  • question1.ows: Orange workflow file (open in Orange Data Mining, free and open source)

Why Orange

Orange was used here instead of a Python script because the task called for a visual, auditable pipeline where every preprocessing and modelling step is inspectable as a widget graph. It's a legitimate tool for data mining coursework and rapid EDA, and this project uses it for genuinely non-trivial work: Welch's t-tests, SHAP explanations, and Bayesian model comparison, not just drag-and-drop basics.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors