Accurate on average isn't enough
Back from my travels with a travel-related machine-learning failure:
Credit card companies use machine-learning techniques to spot suspicious activity: If I live in Vancouver and suddenly my credit card gets charged to buy a diamond necklace in Edinburgh, something's off.
But in my case, despite putting in a travel notice ("I'll be in Poland and Germany"), my credit card was flagged and temporarily disabled because I used it to rent a car at the airport. You know, one of the very first things someone does in a foreign country after landing there.
I'm sure the credit card company did rigorous validating of their fraud detection system. And I'm sure in aggregate it shows sufficiently good results. But aggregate and average are no solace for me when I have to spend two hours on the phone getting my credit card unlocked for a transaction that shouldn't have raised any eyebrows.
The lesson for anyone deploying machine learning solutions: Don't just rely on overall metrics. Come up with specific, high-stakes scenarios and see how the solution performs in them. Product-centric thinking all the way.
