End-to-end data mining analysis of a 1.66M-row US Census dataset, built in Orange Data Mining. Covers preprocessing, exploratory analysis, hypothesis testing, and classification, with a focus on identifying which factors actually predict income.
Tools: Orange Data Mining, SHAP, Welch's t-test, Gradient Boosting, Bayesian model comparison
Starting from a 1.66 million row census file, this project builds a full pipeline: sampling, cleaning, feature engineering, exploratory analysis, statistical testing, and predictive modelling, then extends into a geo-spatial analysis linking income and education to 2020 US election results.
The workflow is built entirely in Orange's visual programming environment rather than a Python/pandas script. The full write-up, including all figures and reasoning, is in report.pdf. This README summarises the approach and findings.
1. Preprocessing
- Randomly sampled 5,000 of 1.66M rows to reduce computation while preserving population distribution
- Converted
statefrom numeric to categorical via a lookup table - Recoded
sex,race,place of birth,class of worker, andmarital statusinto simplified categories using the Formula widget - Grouped
educationinto no diploma / high school / post-high-school / degree bands, andoccupationinto SOC-based industry groups - Removed outliers (415 rows, 8.3%), primarily extreme income (>$500K) and hours (>80/week) values, reducing mean income skew from $56,108 to $54,785
2. Exploratory analysis
- Confirmed strong right-skew in income (typical of income data), corrected with a
log(income + 1)transform - Verified the skew is Pareto-like: a log(rank) vs log(income) plot showed slope ≈ -1.0, meaning the top 1% of earners hold a disproportionate share of total income
- Ran Welch's t-tests across sex, race, and place of birth:
- Sex: $22,743 gap (t = 11.72, p < 0.0001)
- Race: $8,351 gap (t = 3.74, p < 0.0002)
- Place of birth: no significant effect (t ≈ 1.0, p ≈ 0.3)
3. Predictive modelling
- Split income at the median ($38,010) into low/high subgroups
- Compared 7 classifiers; Gradient Boosting performed best (AUC 0.879, accuracy 0.787), confirmed via Bayesian pairwise comparison (0.875 probability of outperforming Random Forest)
- Used SHAP values and permutation importance to explain predictions: hours worked, occupation, education, and age were the strongest predictors, education included
Key finding: although education correlates with income on its own (r = 0.67), it drops to a secondary factor once hours worked and occupation are accounted for. Time on the job and industry sector matter more than years of schooling.
4. Geo-spatial extension
- Joined census income/education data with 2020 election results by state
- States with higher income and education aligned closely with Biden-voting states
- Education showed a significant link to voting outcome (t = 2.88, p < 0.01); income alone did not (t = -1.92, p ≈ 0.10)
- Split states into coastal/inland: income predicted Democratic vote share more strongly in coastal states (r = 0.44) than inland (r = -0.01), though the difference fell just short of significance (Fisher z = 1.598, p = 0.055)
- 5,000-row sample may under-represent rare subgroups and extreme high earners present in the full 1.66M-row dataset
- Binary groupings (white/non-white, male/female, coastal/inland) simplify analysis but remove within-group variation and intersectional effects
- Aggregate state-level correlations do not imply individual-level causation
- DC and Puerto Rico were excluded due to data alignment issues; DC's exclusion likely weakens the income/education-voting correlations, since it is a high-income, high-education, heavily Democratic outlier
- "Coastal" as a binary category is a rough proxy that likely captures urbanicity more than geography itself
report.pdf: full write-up with all figures, workflows, and statistical detailquestion1.ows: Orange workflow file (open in Orange Data Mining, free and open source)
Orange was used here instead of a Python script because the task called for a visual, auditable pipeline where every preprocessing and modelling step is inspectable as a widget graph. It's a legitimate tool for data mining coursework and rapid EDA, and this project uses it for genuinely non-trivial work: Welch's t-tests, SHAP explanations, and Bayesian model comparison, not just drag-and-drop basics.