Fraud Detection on Highly Imbalanced Transaction Data
Investigation at a glance
- Transactions
- 284,807
- Fraud cases
- 492
- Fraud rate
- 0.172%
- Imbalance
- 578:1
- Model
- XGBoost
- Selected threshold
- 0.75
- Recall
- 76%
- False positives
- 20
- False negatives
- 18
Signal
Credit card fraud detection on the Kaggle creditcardfraud dataset: 284,807 transactions, only 492 fraudulent (0.172%).
Standard accuracy is meaningless here: predicting "legit" for every transaction scores 99.83%. The real challenge is catching fraud without drowning analysts in false alarms.
Investigation
Approach
- Stratified 70/15/15 train/validation/test split, preserving the fraud ratio in each set
- Exploratory analysis on transaction Amount: found fraud's median ($9.25) is lower than legit's ($22), despite a higher mean, revealing heavy right-skew
- Baseline logistic regression (scaled features, fit on train only)
- XGBoost with
scale_pos_weightset to the class ratio (578:1) to counter the imbalance - Threshold sweep across 101 cutoffs (0.00 → 1.00), evaluating precision, recall, and error counts at each
Results
| Model | AUC | Precision | Recall | F1 |
|---|---|---|---|---|
| Logistic Regression | 0.9571 | 0.786 | 0.595 | 0.677 |
| XGBoost | 0.9765 | 0.571 | 0.811 | 0.670 |
XGBoost won on AUC and recall, catching 60 of 74 frauds vs 44, but at the cost of more false alarms (45 vs 12). F1 stayed flat because precision and recall moved in opposite directions.
Decision
Rather than accept the default 0.5 threshold, I swept all cutoffs and compared three operating points.
| Operating point | Threshold | Precision | Recall | FP | FN |
|---|---|---|---|---|---|
| Max F1 | 0.96 | 0.92 | 0.73 | 5 | 20 |
| ChosenSelected | 0.75 | 0.74 | 0.76 | 20 | 18 |
| ≥90% recall | 0.01 | 0.03 | 0.92 | 2,148 | 6 |
Chose t=0.75, a recall-weighted threshold capped at 20 alerts, a workload a real team can process. Relative to the default, this cut false alarms from 45 to 20 at the cost of 4 additional missed frauds.
Key takeaway
The model sets the ceiling; the threshold sets the outcome.
Same model, three cutoffs, and false alarms swung from 5 to 2,148, a 400× range driven entirely by a business decision, not a modeling one.
- Python
- pandas
- scikit-learn
- XGBoost
- matplotlib