Skip to Main Content
Back to Insights

What enterprise retailers get wrong when they evaluate AI

What enterprise retailers get wrong when they evaluate AI

Why the vendor demoing the flashiest agent rarely moves your return rate

A doctor who correctly reads the scan but labels the condition wrong doesn’t help the patient. The treatment that follows is built on a real result and the wrong conclusion.

This plays the same out across every stage of the enterprise AI buying cycle. Real tech, genuine deployment – wrong classification.

Buyers who think they’re deploying AI are often deploying a chat interface on top of incomplete data. Buyers who think they’re avoiding risky AI are blocking the exact models most capable of protecting their margin. On top of that, the finance leaders now being asked to approve AI budget lines are walking into those conversations without the vocabulary to tell the difference.

Most enterprise retailers have already funded at least one AI vendor that didn’t move a number.

Being at the front of enterprise buyers at every stage of their AI journey — from first demo to annual quarterly business review — Vishal Patel’s observation about where these conversations go wrong is that most organizations are letting the wrong question drive the budget.

As Chief Product and AI Officer at Appriss Retail, Patel is using his expertise and PhD in computer science to teach deep machine learning at the university level across Appriss where AI-powered systems process decisions across 40% of all U.S. retail transactions. 

Key Takeaways:

  • Single-channel AI misidentifies fraud and abuse. A customer returning across channels looks like multiple low-risk people in a siloed system.
  • Every AI pilot that fails to define financial success before deployment ends the same way: with a demo the organization can’t convert to a board-level number.
  • The AI compliance review is now standard at enterprise retailers. Vendors who can explain exactly which layer makes which decision can pass it.

A budget misdiagnosis

The AI budget conversation at most enterprise retailers centers on: 

  • Are we using it? 
  • Is it approved? 
  • Is it generative? (raise your hand if you can genuinely define what that means)

Those questions made sense when AI was new, but they don’t predict financial return now.

Loss prevention and returns management vendors learned fast that “AI-powered” moves deals. What that label often describes is a chat interface on incomplete data that still hands every decision to a human. These systems produce risk scores, surface alerts, and summarize findings, but they don’t tell a retailer, at the moment a return hits the POS, whether to approve, warn, or decline.

As Patel puts it: “Enterprises don’t always understand this because of the significant amount of AI washing happening in the industry.” For finance leaders approving AI line items, this is a due diligence risk with a financial consequence.

A system that produces risk scores and waits for a human is the equivalent of a test that returns results and waits for the patient to treat themselves. Funded systems that require human judgment at every decision point report on return fraud and abuse, but they don’t reduce it.

decisioning AI vs recommendation AI

The hallucination concern in the wrong layer

Skepticism about generative AI in enterprise environments is legitimate, considering that large language models can produce inaccurate outputs. (Remember when we all discovered ChatGPT is horrible at math).

The problem is that most enterprise buyers are applying that skepticism to the entire AI category, including machine learning models that don’t share the property they’re worried about.

“How many times do humans get confused? How many times do you have an employee who you’ve given a ton of information, and they’re making mistakes?” asks Patel. “Hallucination is essentially an inaccurate result coming out of these types of models.”

Appriss Retail’s AI stack runs on three layers

  • Layer 1 links customer identity across in-store, online, and customer service channels into a single behavioral record. 
  • Layer 2 is where ML models run — scoring that unified record against tagged fraud and abuse patterns to produce an approve, warn, or decline decision. 
  • Layer 3 is where investigators work: dashboards, alerts, case summaries.

The hallucination concern that’s circulating in enterprise AI conversations belongs to a different layer entirely. ML models are deterministic. Given the same behavioral pattern on a shopper, they produce the same output every time. There is no generative step, no hallucination risk.

“Hallucination is not a concept that exists in machine learning models,” explains Patel. “Given the same behavior pattern about a shopper, they’re going to make the same recommendation. Something auditable, compliant, and deterministic.”

The practical consequence of misapplying this concern: retailers who block AI from their return authorization workflows to avoid generative AI risk are blocking the only model type that doesn’t carry that risk. Ecommerce is a channel that is often excluded. Total U.S. retail returns reached over $700 billion last year, with roughly $100 billion estimated as fraudulent or abusive. A system that can only score in-store transactions is structurally incapable of moving that number.

3 criteria that predict whether AI produces ROI

Patel organizes AI ROI around three criteria of what a correct diagnosis actually requires: complete patient history, the right test for the right condition, and a defined recovery metric before treatment begins.

  1. Data foundation before model deployment

The first question to ask any AI vendor: does this system make decisions on a complete picture of the customer (think across in-store, online, and customer service center) or a partial one?

A single-channel view will miss behavior and produce wrong decisions. A customer who buys online and returns in-store across 3 different locations looks like 3 separate low-risk people in a siloed system. “Most of the other players in the space are purely focused on a single channel, creating a partial view of the shopper,” says Patel. Appriss Retail’s consortium covers 2.5B+ transactions every month.

Any vendor can license a model, but 20+ years of tagged fraud and abuse behavior across wardrobing, receipt reuse, and tender laundering patterns at scale is not replicable.

  1. Determinism at the decision layer

The right model for a real-time return authorization is a supervised ML model trained on labeled fraud data, with known inputs and reproducible outputs. The model isn’t making critical approve/warn/decline calls, the machine learning layer does that. 

Generative AI does belong in the workflow layer to accelerate investigation and summarize cases, but not at the decision point. A vendor who can’t tell you which layer is making which call will struggle when IT and audit ask the same question.

  1. Define financial success before deployment

Most enterprise AI pilots fail to define what success looks like before the contract is signed. The CFO-ready version of this conversation is simple: the industry average return rate was 14.2% in 2025, against $706 billion in total returns. If that number moves two percentage points, the dollars saved are what get taken to the board.” Patel makes it clear:

“This is not a prototype or a pilot AI project you’re unsure of the ROI on. This is an AI investment with a well-defined outcome.” 

A vendor who can’t frame it that way isn’t ready for a CFO conversation.

Ask your vendor: identity resolution, decision vs analytics and data foundation

The next wave of evaluation

The next shift is heading towards agentic AI that is generative; systems that act inside workflows without waiting for a human to initiate them. Sidekick, Appriss’s AI collaborator embedded across Secure, Engage, and Incident, already writes searches, analyzes exceptions and returns, and generates case summaries inside the platform.

What’s top of mind now is an overnight processing layer where agents score exceptions and return trends, so investigators log in to a prioritized queue and AI-validated insights instead of raw data.

“The actions are still human-led, and are always going to be human-led.” 

The agentic roadmap doesn’t change which decisions humans make, rather how much time gets consumed before they make them.

The AI compliance review is now standard at enterprise retailers: which layer makes which decision, where the data goes, what happens when the model is wrong. If you’ve built on deterministic ML and proprietary data, you can answer those questions. Vendors who built a chat interface on a third-party LLM aren’t misdiagnosing the problem. They’re misrepresenting what they treat.

Frequently asked questions

What’s the difference between AI that makes decisions and AI that makes recommendations?

Decisioning AI produces an approve, warn, or decline output at the moment of a transaction — no human required. Recommendation AI surfaces a risk score that an associate then acts on. In high-volume, omnichannel return environments, that gap compounds fast. Associates miss alerts, inconsistently apply risk scores, and make different calls on the same signals. A decisioning system doesn’t.

How do I evaluate AI vendor claims when every vendor says they have AI?

Three questions cut through: Which layer of your stack makes the risk decision — ML model, rules engine, or LLM? What does your training data actually contain — is it tagged for wardrobing, tender swapping, tender laundering specifically? Can you show me pre- and post-deployment return rates for a comparable retailer?

Vendors who answer all three specifically are worth the next conversation. Vendors who lead with model count or years of AI investment usually can’t answer the third.

Why can’t I use a single-channel system for return fraud detection?

A single-channel system produces wrong decisions. A customer returning across three channels appears as three separate low-risk individuals in a siloed system. Unified identity across channels exposes the pattern immediately. With 29% of all returns now BORIS and nearly half of customers using omnichannel return methods, a system that only scores one channel is operating on a partial picture of your highest-risk transactions.