Loading prices…
STKR NewsSTKR News0 of 3 free this month
AI

Newer AI models missed more payment fraud in Coinbase’s benchmark

New research from Coinbase suggests that bigger isn't always better for AI, as newer models actually missed more payment fraud in recent testing.

Originally on CryptoSlate →
AB

Adrian Boysel

Contributor

Oct 10, 2026

5 min read

Photo illustration / STKR News

In the world of building crypto products, we are often told that the newest iteration of a large language model is the definitive solution to our problems. We see the benchmarks for coding, creative writing, and logic, and we assume those gains translate to the gritty, high-stakes work of financial security. But a recent deep dive from the team at Coinbase serves as a sobering reality check for anyone building on the frontier of AI and payments.

The Shiny Object Trap

Coinbase recently conducted a fixed historical replay to see how different generations of AI models handle payment fraud. For those of us in the trenches, this is the ultimate test. You aren't asking the AI to write a haiku; you are asking it to stop a thief from draining a user's wallet. The results were not what the marketing departments at major AI labs would have you believe. In their internal benchmarking, newer model upgrades actually demonstrated weaker fraud coverage than their predecessors.

This is a massive red flag for founders. We have been conditioned to believe in a linear path of progress: GPT-4 must be better than 3.5, and 4o must be better than 4. But when it comes to the specific, adversarial environment of fraud detection, the data suggests that these general-purpose improvements might be diluting the specific logic needed to catch bad actors.

Understanding Precision versus Recall

To understand why this matters, we have to look at how these models are measured. Coinbase tracked two main metrics: precision and coverage (or recall). Precision tells you how often the model is right when it screams "fraud." Coverage tells you how many total frauds it actually caught. If your precision is high but your coverage is low, you are catching the obvious thieves but letting the sophisticated ones walk through the front door.

The Coinbase benchmark found that while precision improved in certain GPT iterations, the overall ability to cast a wide net and capture fraud dropped in three consecutive model upgrades. For a builder, this creates a dangerous trade-off. Do you want a model that is polite and rarely wrong, or do you want a model that is actually effective at securing your treasury?

The Founder's Dilemma

If you are building a dApp or a fintech bridge, your instinct is to integrate the latest API the moment it drops. We want to tell our investors we are using the cutting edge. But this data suggests that "upgrading" your security layer could actually be a downgrade in protection. It highlights a fundamental truth about AI in 2024: these models are becoming more conversational and more "human-aligned," but that alignment often comes at the cost of the raw, clinical pattern matching required for risk management.

Why Newer Isn't Always Better

There are a few reasons why we might be seeing this regression in fraud detection capabilities:

  • RLHF Over-Optimization: Reinforcement Learning from Human Feedback makes models easier to talk to, but it can also make them more hesitant to flag edge cases as "bad" if they don't look like typical examples of malice.
  • Generalization Loss: As models try to be good at everything—from medical advice to Python coding—they may lose the specialized "neurons" that were previously very good at spotting financial anomalies.
  • Data Drift: The way people commit fraud evolves faster than the training sets for these massive models. If a model is trained on a broader set of general data, it might lose its edge on the specific, fast-moving signals of crypto-native fraud.

What This Means for Builders

We need to stop treating AI models like magic black boxes that only get smarter. If you are a founder or an engineer responsible for security, the Coinbase findings should change your workflow. You cannot trust the benchmark reports coming out of OpenAI or Anthropic when it applies to your specific risk profile.

First, you need your own historical replay. You should have a vault of every fraudulent transaction your platform has ever seen. When a new model comes out, you don't switch over immediately. You run that new model against your vault and see if it catches what the old one caught. If it doesn't, you don't upgrade—period.

Second, precision is a vanity metric if your coverage is failing. It’s easy to be right if you only guess when you’re 100% sure. But in the world of crypto payments, the 20% of cases where the model is "unsure" are often where the real damage happens. You need a system that prioritizes coverage, even if it requires a human-in-the-loop to handle the slight increase in false positives.

The biggest risk in crypto isn't the technology failing; it's the assumption that the technology is getting better at protecting us just because it's getting better at talking to us.

The Skeptic's Path Forward

I’ve always said that builders in this space need to be the biggest skeptics of their own stacks. The Coinbase report is a gift because it provides the data to back up that skepticism. It shows that the "AI revolution" is not a rising tide that lifts all boats equally. Some boats—specifically those tied to security and fraud detection—might actually be taking on water as the models become more "refined."

For those building the next generation of financial tools, the takeaway is clear: audit your AI as strictly as you audit your smart contracts. Don't let a version number fool you into thinking your users are safer. In the game of cat and mouse between developers and fraudsters, the mouse just got a little bit smarter, and the cat just got a little bit more distracted by trying to be a better conversationalist.

Practical Steps for Your Team

  • Version Pinning: Don't use "latest" tags in your production environment. Pin your models to specific versions that have been verified against your own fraud datasets.
  • Hybrid Logic: Never rely solely on an LLM for fraud. Use hard-coded heuristics and traditional machine learning models alongside AI to catch what the newer models are missing.
  • Transparency: Be honest with your users about how you are using AI for security. If the tech is regressing, you owe it to them to have secondary layers of defense in place.

We are in an era where the hype around AI is blinding us to the technical regressions happening under the hood. Coinbase’s willingness to publish these findings is a win for the builder community. It reminds us that our job isn't to use the coolest tech—it's to build things that actually work and keep people's money safe. Sometimes, that means sticking with the old guard while the new models find their footing.


Read the original at CryptoSlate →

The Brief

Stay Updated on Cutting-Edge Tech

A six-minute morning dispatch on the markets and the technology shaping them.

Free. No spam. Unsubscribe anytime.

Write for STKR

Become a Contributor

Earn $STKR for published stories on markets, protocols, and culture.

  • Earn $STKR for every published piece
  • Editorial support from the STKR desk
  • Byline visibility across the network
  • First look at the upcoming creator program
Apply to Write

Keep reading

All stories

Comments

24 reader responses