Data AnalystFraud & Risk AnalyticsFintech

I find where the money moves wrong.

I analyze transaction and user data to trace where fraud, risk, and revenue actually move, then build the pipelines, models, and dashboards that let risk and product teams act on it.

Nairobi, Kenya

Open to Data Analyst, Fraud/Risk Analyst, Credit Risk Analyst, and Junior Data Scientist roles

01 / Profile

About

BSc Statistics graduate who worked inside a live fintech payments environment, where data quality, user behaviour, and fraud signals meet.

I'm looking for Data Analyst, Fraud/Risk Analyst, Credit Risk Analyst, or Junior Data Scientist roles, particularly in fintech and financial services.

Experience: Hurupay, fintech internship

  • AnalysisBuilt anomaly detection scripts directly on payment event data.
  • DatabasesWrote multi-table SQL across MongoDB and PostgreSQL to support fraud workflows.
  • IntegrationPulled data from open banking APIs to enrich internal datasets.
  • ContextPaired the raw numbers with behavioral context from Amplitude and Intercom, because a flagged transaction and a confused user are often the same event seen from two sides.
  • DeliveryWorks in Python, R, and SQL, and builds dashboards people without a data background can actually use.
50+
students supported as Data Science TA, JKUAT
15%
average project score improvement under that mentoring
2
production databases worked with: MongoDB & PostgreSQL
2026
BSc Statistics · JKUAT · 2026

How I work

  1. 01

    Data

    Payment event data and open banking API pulls, held in MongoDB and PostgreSQL.

  2. 02

    Context

    Behavioral context from Amplitude and Intercom alongside the raw numbers.

  3. 03

    Analysis

    Python, R, and SQL, including anomaly detection scripts.

  4. 04

    Decision

    Dashboards that people without a data background can use.

02 / Work

Case Files

Selected work, ordered by relevance, not by date.

Case 01 / Featured investigationStatus: Threshold tuned

Fraud Detection on Highly Imbalanced Transaction Data

Investigation at a glance

Transactions
284,807
Fraud cases
492
Fraud rate
0.172%
Imbalance
578:1
Model
XGBoost
Selected threshold
0.75
Recall
76%
False positives
20
False negatives
18
01

Signal

Credit card fraud detection on the Kaggle creditcardfraud dataset: 284,807 transactions, only 492 fraudulent (0.172%).

Standard accuracy is meaningless here: predicting "legit" for every transaction scores 99.83%. The real challenge is catching fraud without drowning analysts in false alarms.

02

Investigation

Approach

  • Stratified 70/15/15 train/validation/test split, preserving the fraud ratio in each set
  • Exploratory analysis on transaction Amount: found fraud's median ($9.25) is lower than legit's ($22), despite a higher mean, revealing heavy right-skew
  • Baseline logistic regression (scaled features, fit on train only)
  • XGBoost with scale_pos_weight set to the class ratio (578:1) to counter the imbalance
  • Threshold sweep across 101 cutoffs (0.00 → 1.00), evaluating precision, recall, and error counts at each
03

Results

Model comparison
ModelAUCPrecisionRecallF1
Logistic Regression0.95710.7860.5950.677
XGBoost0.97650.5710.8110.670

XGBoost won on AUC and recall, catching 60 of 74 frauds vs 44, but at the cost of more false alarms (45 vs 12). F1 stayed flat because precision and recall moved in opposite directions.

04

Decision

Rather than accept the default 0.5 threshold, I swept all cutoffs and compared three operating points.

Operating points
Operating pointThresholdPrecisionRecallFPFN
Max F10.960.920.73520
ChosenSelected0.750.740.762018
≥90% recall0.010.030.922,1486

Chose t=0.75, a recall-weighted threshold capped at 20 alerts, a workload a real team can process. Relative to the default, this cut false alarms from 45 to 20 at the cost of 4 additional missed frauds.

Key takeaway

The model sets the ceiling; the threshold sets the outcome.

Same model, three cutoffs, and false alarms swung from 5 to 2,148, a 400× range driven entirely by a business decision, not a modeling one.

  • Python
  • pandas
  • scikit-learn
  • XGBoost
  • matplotlib
View repository ↗ for Fraud Detection (opens in new tab)

Selected analytics work

Case 02 to 03
Case 02Pattern found

New York Airbnb Listings, 2024 (EDA)

Problem

Surface pricing and demand patterns across 20,000+ NYC listings for a host or platform deciding how to price. Cleaned raw listing data (missing values, type corrections) across boroughs, room types, and review volume.

Finding

Review count correlates positively with price, meaning reputation itself carries measurable pricing power, independent of location or room type.

Tools
  • Python
  • Pandas
  • NumPy
  • Matplotlib
View on GitHub ↗ for New York Airbnb Listings (opens in new tab)
Case 03Excel dashboard

Sales Performance Dashboard

Problem

Give non-technical stakeholders a way to drill into sales by product, region, and time, without touching a formula. Built with Pivot Tables, Pivot Charts, and Slicers for dynamic filtering across every dimension of the dataset.

Takeaway

Designed so a stakeholder could ask a new question of the data live, in the meeting, instead of requesting a follow-up report.

Tools
  • Excel
  • Pivot Tables
  • Power Query
View on GitHub ↗ for Sales Performance Dashboard (opens in new tab)

Adjacent work

Case 04
Case 04Contributor, web app

HealthPoint Kenya

Problem

Help people describe symptoms in plain language and get routed to the right kind of healthcare facility. Contributed to an AI-powered web app; integrated a Google Genkit AI flow to handle natural-language symptom queries, applied AI work outside a pure data-analysis context.

Takeaway

Proof that the same query-and-pipeline thinking behind fraud detection transfers cleanly to a different domain entirely.

Tools
  • Next.js
  • TypeScript
  • Google Genkit AI
View on GitHub ↗ for HealthPoint Kenya (opens in new tab)

03 / Skills

Capabilities

The tools and analytical methods I use.

Analysis

Exploring, testing and summarising data.

  • Python (Pandas, NumPy, Matplotlib, Seaborn)
  • R (dplyr, ggplot2)
  • SQL
  • EDA
  • Statistical analysis
  • Hypothesis testing
  • Time series
  • Feature engineering

Modeling

Predictive and anomaly models.

  • Classification
  • Regression
  • Anomaly detection
  • scikit-learn
  • XGBoost
  • caret

Data

Querying and joining production data.

  • PostgreSQL
  • MongoDB
  • Query optimization
  • Window functions
  • Multi-table joins

Decision support

Reporting and behavioral analytics.

  • Tableau
  • Amplitude
  • Excel (Pivot Tables, Power Query, Slicers)
  • Google Sheets
  • Jupyter

Tools

  • Git
  • GitHub
  • VS Code
  • RStudio
  • Intercom
  • TypeScript

04 / Contact

Looking for the next signal.

Open to Data Analyst, Fraud/Risk Analyst, Credit Risk Analyst, and Junior Data Scientist roles.