Fraud-model accuracy hides the decision that actually costs money
A Red Hat engineering demo shows why teams should rank imbalanced classifiers by business loss before scaling the workflow with OpenShift AI AutoML.
A Red Hat Developer walkthrough makes a useful point for teams moving tabular models into production: model selection is a business decision disguised as a leaderboard. Its synthetic fraud demo gives a naive classifier 95.8% test accuracy even though it catches none of 19 fraudulent transactions. A cost-ranked candidate reaches 97.9% accuracy, catches 15 of 19 and finishes an estimated $3,578 ahead over the 480-transaction test set.
Those dollar figures belong to a demonstration dataset, not a production claim. The reusable part is the selection method.
Rank the failure, not the average
Fraud makes up about 4% of the demo’s 2,400 generated transactions. At that imbalance, a model can look accurate by predicting that nearly everything is legitimate. Recall exposes missed fraud, but optimizing recall alone can send every transaction to human review.
The walkthrough therefore scores each candidate by combining the transaction value recovered when fraud is caught with a fixed $6 review cost for every alert. Area under the ROC curve breaks ties, while accuracy, precision and recall remain visible for diagnosis. The full sweep tests 67 configurations across logistic regression, decision trees and random forests, along with feature sets, fraud-row weights and decision thresholds.
The winner is still logistic regression at a 0.5 threshold. The difference is that it uses all 12 features and gives fraudulent rows four times the training weight. That is a helpful result: better model selection does not necessarily mean a more complex algorithm. It means testing the configuration against the cost of its mistakes.
Keep the test split sealed
The demo uses a stratified 60/20/20 split. Candidates are trained and ranked against validation data; the test set stays untouched until the final comparison. Platform teams should preserve that boundary when translating the pattern into pipelines. A custom business metric becomes unreliable if repeated tuning leaks information from the final test set into the search.
The local implementation deliberately uses explicit Python rather than an automated search service. That makes the cost function and each model choice inspectable before the workflow is moved onto a platform.
Where OpenShift AI fits
The article maps the laptop version onto Red Hat OpenShift AI: object storage replaces the local CSV, AutoML runs an AutoGluon search through Kubeflow Pipelines, the dashboard presents the leaderboard and the selected model can move into the model registry and serving endpoint. AutoML is currently a Technology Preview.
There is an important limitation. For binary classification, the platform feature optimizes accuracy, so a team with a business-specific loss function would apply its own cost calculation to the resulting leaderboard rather than assuming the default winner is deployable.
The demo places an LLM after the classifier. The predictive model decides whether to flag a transaction and supplies evidence; a locally served model drafts the analyst note. In a platform deployment, that second component can be a vLLM-served model. Keeping those roles separate gives operators two measurable controls: the classifier’s decision economics and the language model’s explanation quality.
The practical first step is small: write down the cost of a false negative and a false positive, encode that calculation, and see whether it changes which existing model would ship.
sources
- How to rank fraud detection models using custom cost metricsdevelopers.redhat.com
- OpenShift AI AutoML and AutoRAG announcementwww.redhat.com
- Fraud-model search demo repositorygithub.com
comments · 0