Researchers measured why enterprise AI stalls by testing six regulated financial workflows across 72 model and tool configurations. While 57 of 72 passed a simple demonstration benchmark, only 32 met strict production requirements for sustained accuracy, reproducibility, and verifiable attribution. The study highlights a significant gap between demo viability and operational reliability in high-stakes environments.
- Demonstration success does not predict production readiness; 21 configs failed the production bar.
- Production bars require sustained accuracy, reproducibility, and verifiable attribution, not just single-case correctness.
- Regulated financial services face a 44% failure rate when moving AI from demo to production.
- Confidence signals must carry actual information to meet the new production standards.