Exchanges· ★★★· neutral·

Coinbase benchmark: newer AI models missed more payment fraud

  • —Sonnet recall fell 22.2 percentage points; dollar-weighted recall dropped 22.9 points
  • —Opus recall declined 0.8 points and precision also fell
  • —GPT precision rose 11.5 points, but recall fell 20.7 points
  • —A post-trained Qwen3.5-9B beat Opus 4.5: F1 up 9.6 points, dollar-weighted recall up 35.4 points
Why it matters: The findings challenge the assumption that upgrading a model improves an existing payment screener, pushing providers to test candidates under their own decision setup.
Source: CryptoSlate