Every card transaction streams through Kafka, gets point-in-time features from a stateful engine, is scored by a gradient-boosted model in a fraction of a millisecond, and is approved, sent for review or declined with reasons an analyst can read. A drift monitor watches the traffic; when it shifts, a challenger model is retrained and has to beat the champion before it is promoted.
Everything on this page is a replay of a real run on a laptop: Kafka 4.3, 956k transactions, 15 unseen "live" days. The transactions are synthetic (real fraud data is private); how the data is generated.
Days 60–74 were never seen in training. From day 67 a seasonal shift begins (bigger baskets, more electronics and online shopping), the kind of week that floods fraud teams with false alerts.
The model ranks every transaction; the business decides how many an analyst team can look at. Move the budget and see what it buys on the held-out test period.
The monitor compares live feature distributions with training (Population Stability Index; above 0.25 is drift). It needs no labels, so it fires days before chargebacks would reveal a problem.
Each alert carries the model's top reasons (per-transaction SHAP contributions translated into plain language) and which legacy rules would have fired. Truth is revealed here because this is a replay; the list is a sample of the run's alerts plus every missed fraud.
A seeded generator creates 6,000 cards across 12 Saudi cities and 900 merchants. Legitimate customers have homes, habits, favourite merchants, devices, trips and occasional splurges. Five fraud campaigns are mixed in: card testing, account takeover, stolen cards used far away, a merchant fraud ring on shared devices, and "low and slow" fraud that looks normal (the model catches about half of it, which is the honest hard case). Training uses days 7–41, validation 42–48, test 49–59; the live replay is days 60–74.