• 4 mins read
  • Published

Coinbase Flags AI Model Upgrades for Missing More Fraud

Catheryne Nicholson Crypto infrastructure writer EgonCoin

Post by Catheryne Nicholson

Coinbase Flags AI Model Upgrades for Missing More Fraud EgonCoin © egoncoin.com
Coinbase Flags AI Model Upgrades for Missing More Fraud © egoncoin.com

Coinbase's internal replay found that newer AI models let more fraudulent payments slip through and recovered less fraud value, casting doubt on the payoff of routine model upgrades for payment screening.

Coinbase's engineers ran a historical replay of 16,140 Onramp transactions, pulling data from 7,293 users and 813 confirmed fraud cases. The test spanned nine weeks before the company's risk agent went live. Each AI model candidate faced the same decision policies and guidance, stripping out any changes to the broader screening system. The team pitted Opus 4.5 against Opus 5, Sonnet 4.6 against Sonnet 5, and GPT-5.4 against GPT-5.6 (sol), tracking recall, precision, and dollar-weighted recall to see which models actually caught more fraud and how much value they protected.

In Coinbase's benchmark, Sonnet 5's recall rate dropped by 22.2 percentage points compared to Sonnet 4.6, while dollar-weighted recall fell by 22.9 points.

Analyst

The numbers cut through the hype. Every newer model version posted lower recall, lower F1 scores, and lower dollar-weighted recall. Sonnet's recall rate fell by 22.2 points, with dollar-weighted recall down 22.9. Opus's recall slipped by 0.8 points. Both newer models also lost ground on precision. GPT's latest version was a mixed bag: precision jumped by 11.5 points, but recall and dollar-weighted recall dropped by 20.7 and 21.8 points. The upshot: the model flagged fraud more accurately, but let more bad payments through.

Coinbase's replay didn't track actual customer losses from these newer models or pinpoint why the regressions happened. The company stressed that the evaluation isolated model performance under a fixed policy, not a redesigned screening stack. That's a warning shot for payment providers. Swapping in a new AI model without tuning the entire decision setup can backfire, leaving fraud coverage weaker than before.

In a separate run, Coinbase post-trained a Qwen3.5-9B model using historical fraud outcomes and deterministic rewards. This model beat Opus 4.5 on four fraud-detection metrics. F1 improved by 9.6 points and dollar-weighted recall surged 35.4 points. Median end-to-end request latency dropped to 0.683 seconds, a 55% cut from Opus 4.5's 1.515 seconds. These results came from a different evaluation and don't resolve the main upgrade shortfalls.

Coinbase Onramp combines traditional machine learning and rule engines with selective review by a large language model, analyzing user behavior and recent transactions. This hybrid approach supports crypto purchases through partner apps, including guest checkout modes.

TokenPost

For payment and crypto platforms, the findings cut against the idea that AI model upgrades are a plug-and-play fix for fraud. Coinbase now tells companies to test candidate models inside their actual decision setup before changing prompts, thresholds, or operational parameters. Latency and reliability need to be weighed alongside detection quality. The lesson echoes other crypto infrastructure headaches, where technical upgrades sometimes bring new trade-offs or side effects. As reported earlier, even well-meaning improvements can open the door to fresh problems.

Coinbase's historical replay covered 16,140 transactions and 813 confirmed fraud cases over nine weeks before its risk agent arrived. The company compared multiple versions of Opus, Sonnet, and GPT model families, finding every newer version underperformed its predecessor in recall, F1, and dollar-weighted recall. In contrast, a post-trained Qwen3.5-9B model delivered a 9.6-point F1 boost and a 35.4-point jump in dollar-weighted recall, while halving median request latency compared to Opus 4.5.

Fraud detection in crypto payments hinges on how models are integrated and tuned within real-world decision setups. High recall means more fraud gets caught, but if precision drops, more legitimate transactions get flagged, driving up user complaints and operational costs. Dollar-weighted recall tracks how much fraud value is actually stopped. Payment providers have to balance these metrics, since chasing one can erode another. Coinbase's results show that model upgrades alone don't guarantee better fraud prevention. The real gains come from careful integration and context-specific testing inside the broader payment stack.

Related articles