A till receipt where every item reads found except one marked wrong item, wrong charge. Each item is found 96.45 to 97.34% of the time, yet the whole basket is right 42.48%. Beside it, baskets exactly right on the RPC checkout dataset: 73.17% at 3 to 10 items, 54.69% at 10 to 15, 42.48% at 15 to 20.

Computer vision in retail: which uses pay back, and why integration costs more than the cameras

Shelf alerts and queue counts earn their keep in ordinary stores. Checkout-free, the famous one, now lives mostly in small stores, and the reason is arithmetic you can check yourself.

What actually works in computer vision in retail?#

Computer vision in retail works where a missed detection costs little and feeds one task list, and stalls where every item must be right and must reach payment. So every use case on a vendor's list faces two tests. First, what does one wrong detection cost? Second, how many store systems must a right detection reach before it pays?

The two testsTwo tests sort every retail camera use case: what one miss costs, and how many store systems a right answer must reach. Drawn from the RPC paper (22 January 2019) and Amazon's update of 27 January 2026.

A shelf-gap alert passes both tests. When the camera misses one empty facing, the next pass catches it. Also, the alert only has to reach the restocking list. A queue count works the same way, because a count that is off by one changes nothing. Then its output goes to whoever opens a till. In short, these uses forgive error and touch one system.

Checkout-free fails both tests. Every item in the basket must be identified, because one wrong item is a wrong charge. Then the result must reach payment, the point of sale (POS) system and inventory at once. Amazon's own update, revised on 27 January 2026, opens with "Amazon is closing Amazon Go and Amazon Fresh physical stores". In particular, the same article explains why its large grocery stores had already moved to a cart. What follows is the money, the mechanism and the evidence, then three numbers a pilot should prove.

How much money is a retail camera project actually chasing?#

Retailers in the NRF's 2023 survey lost 1.6 percent of sales to shrink, $112.1 billion, up from 1.4 percent and $93.9 billion a year earlier. Shrink is stock that leaves the store unpaid, through theft, error or damage. The National Retail Federation published these figures in its National Retail Security Survey on 26 September 2023.

Shrink loss at US retailers, fiscal 2021 against 2022$93.9 billion against $112.1 billion

$93.9 billion

Fiscal 2021, 1.4 percent of sales

$112.1 billion

Fiscal 2022, 1.6 percent of sales

Shrink rose in both share and dollars between fiscal 2021 and 2022.

Shrink loss at US retailers, fiscal 2021 against 2022 (billions of US dollars of shrink)
Optionbillions of US dollars of shrink
Fiscal 2021, 1.4 percent of sales$93.9 billion
Fiscal 2022, 1.6 percent of sales$112.1 billion

Source: National Retail Federation, National Retail Security Survey, 26 September 2023

However, empty shelves cost more than theft does. Gruen and Corsten wrote a 2007 report for three US grocery industry groups (GMA, FMI and NACDS). In it they restate that "retailers on average lose 4 percent of their annual sales due to OOS items", meaning out-of-stocks. So a shelf camera chases a loss of a few percent of sales, not a fortune.

For example, say a store has 20 million US dollars of annual sales. In that worked example, about 800,000 dollars a year walks out as missed sales at 4 percent. In the same worked example, shrink at the NRF's 1.6 percent adds about 320,000 dollars more. Therefore a camera project there must win back a share of those two figures, after its own running costs.

What empty shelves and shrink cost your store

Enter your store's annual sales; it applies the 4 percent out-of-stock loss and 75 percent store-practice share from Gruen and Corsten, and the NRF's 1.6 percent shrink.

Your store

an assumption, change it

Only out-of-stock loss and shrink are counted. Camera, recognition, catalogue and staff costs are left out because they differ by store and vendor.

Yearly out-of-stock loss at 4 percent of sales

$800,000

Share of it that starts in store practice, at 75 percent
$600,000
Yearly shrink at 1.6 percent of sales
$320,000

Illustrative, modelled not measured; rates from Gruen and Corsten (2007) and the NRF (26 September 2023).

Why do shelf cameras pay back only when the alert reaches a person?#

Gruen and Corsten found that 75 percent of out-of-stocks start in store practice, so a shelf camera earns its cost only when its alert becomes a finished task. Their 2007 report says "75 percent of the cause was due to retail store practices (opposed to up-stream supply issues)". In practice, most gaps come from ordering, shelf filling and stock left in the back room.

Show data table
Three in four out-of-stocks start inside the store: retail store practices 75 percent, upstream supply issues 25 percent. Source: Gruen and Corsten for GMA, FMI and NACDS, 2007.
Segment Value (percent of out-of-stock causes) Share
Retail store practices 75 75%
Upstream supply issues 25 25%

Three in four out-of-stocks start inside the store, so the fix is a person, not a supplier.

Figure Three in four out-of-stocks start inside the store: retail store practices 75 percent, upstream supply issues 25 percent. Source: Gruen and Corsten for GMA, FMI and NACDS, 2007. Gruen and Corsten for GMA, FMI and NACDS, 2007

Seeing the gap is the easier half. Goldman and colleagues built the SKU-110K shelf photo set, released on 1 April 2019, with well over a hundred products packed into each photo. They call precise detection in such packed scenes "a challenging frontier". However, a shelf count does not need perfection, because a missed facing is caught on the next pass.

The hard half is what happens after the alert. Because most gaps start in the store, the fix is a person carrying stock from the back room. So the alert must land in the restocking list that staff already use, with the aisle, shelf and product named. Then someone must close the task, and the closed task must update inventory. Without that loop, a shelf camera makes an accurate list of problems that stay unfixed. In the worked example above, the loop is worth up to 600,000 dollars a year; the camera alone is worth none of it.

Why does a model that reads a shelf well fail a whole basket?#

On the RPC checkout dataset, the best method got the whole basket right 73.17 percent of the time with light clutter and 42.48 percent with heavy clutter. Checkout accuracy is the product of every item being right, so it falls as baskets grow. Wei, Cui, Yang, Wang and Liu published the dataset on arXiv on 22 January 2019. It holds 200 product types photographed on a checkout counter.

The authors scored a basket as passing "if and only if the complete product list is accurately predicted". With 3 to 10 items, the best method passed 73.17 percent of baskets in 2019. Then with 10 to 15 items it passed 54.69 percent, and with 15 to 20 items only 42.48 percent.

Show data table
Whole-basket accuracy falls as baskets grow, even for the best method tested. Source: Wei, Cui, Yang, Wang and Liu, arXiv 1901.07249, 22 January 2019
Item Value
Easy clutter 73.17
Medium clutter 54.69
Hard clutter 42.48

Only 42.48 percent of the largest baskets came out exactly right, against 73.17 percent of the smallest.

Figure Whole-basket accuracy falls as baskets grow, even for the best method tested. Source: Wei, Cui, Yang, Wang and Liu, arXiv 1901.07249, 22 January 2019 Wei, Cui, Yang, Wang, Liu (arXiv 1901.07249), 22 January 2019

Meanwhile, the same 2019 paper reports a standard per-item detection score between 96.45% and 97.34% at every basket size. That gap is the whole story. Each item is found well, but a basket is only right when every item is right, so small errors multiply.

For example, say each item is right 99 times in 100. In that worked example, a five-item basket is fully right about 95 percent of the time. With fifteen items, the same worked example drops to about 86 percent. While a shelf count can live with that, a till cannot, because each failure is a wrong charge to a real customer.

What did Amazon's own stores show about checkout-free?#

Amazon kept Just Walk Out for small stores with quick, few-item trips, moved large grocery stores to Dash Cart, and in January 2026 began closing Amazon Go and Amazon Fresh. Its own Just Walk Out site says the system "uses computer vision, weight sensors, and deep learning models to identify the items". In short, even the flagship does not rely on cameras alone.

Amazon's article on Just Walk Out and Dash Cart explains the split in plain terms. Shoppers in small stores are "making quick purchases of relatively few items". However, in larger grocery stores, "where customers are making a big weekly trip and buy a greater number of items", they "so far prefer Amazon Dash Cart". With Dash Cart, shoppers sign in and start "scanning, and weighing items as they go". Then the update of 27 January 2026 added the closures, with some stores turning into Whole Foods Market stores.

That is the RPC curve playing out in real aisles. Where baskets are small, checkout-free works and sells well. For instance, Amazon reports an 85% increase in transactions per game at Lumen Field in Seattle after its first Just Walk Out store opened. Also, the Just Walk Out FAQ lists stadiums, venue concessions and event merchandise shops. It offers store kits that are "faster to deploy". Where baskets are big, though, the company that built the system chose a cart where the shopper does part of the identifying. That is a design choice worth copying, not a failure to hide.

Amazon matches the checkout design to the size of the trip: cameras alone for few items, a scanning cart for a weekly shop, RFID for apparel. Sources: Amazon, Just Walk Out and Dash Cart article updated 27 January 2026; Just Walk Out FAQ, accessed 1 October 2026.

TripDesign Amazon usesWhat identifies the items
Small store, quick trip with few itemsJust Walk OutComputer vision, weight sensors and deep learning models
Large grocery store, big weekly tripDash CartThe shopper scans and weighs items as they go
ApparelRFID lanesRFID readers at the exit gate read the tags in the clothing

Which retail losses can a camera actually see?#

In the NRF's 2023 survey, external theft was 36 percent of shrink, internal theft 29 percent and process failures 27 percent, and each needs a different camera and workflow. The survey says external theft "accounted for an average of 36% of total loss". Meanwhile, unknown causes made up 6 percent and other causes 1 percent.

Show data table
Shrink has three large sources, and only one of them happens at the shop door. Source: National Retail Federation, National Retail Security Survey, 26 September 2023
Item Value
External theft 36
Internal theft 29
Process, control failures and errors 27
Unknown 6
Other 1

External theft is the largest source at 36 percent, with internal theft and process failures close behind.

Figure Shrink has three large sources, and only one of them happens at the shop door. Source: National Retail Federation, National Retail Security Survey, 26 September 2023 National Retail Federation, 26 September 2023

Each source is seen in a different place. External theft happens on the shop floor and at the exits. So the camera watches aisles and doors, and store security acts on it. Internal theft happens at tills, in stock rooms and at the receiving dock. There the camera watches those spaces, and a loss team acts later. Finally, process failures are wrong counts, wrong labels and damaged stock. Those show up in receiving and in the stock records more than on any video.

Each source of shrink is seen in a different place and acted on by a different team. Source: National Retail Federation, National Retail Security Survey, 26 September 2023.

Source of shrinkShare of shrinkWhere it shows up, and who acts
External theft36 percentThe shop floor and the exits; store security acts
Internal theft29 percentTills, stock rooms and the receiving dock; a loss team acts later
Process failures27 percentReceiving and the stock records, more than on any video

As a result, one camera pointed at all of shrink fails. A project should name one source, one place to watch and one team that acts. Then it can be measured against that slice of the loss, not against the whole 1.6 percent.

What does a loss-prevention camera cost when it is wrong?#

The FTC found Rite Aid's facial recognition produced thousands of false-positive matches between 2012 and 2020, and banned the retailer from using it for five years. The US Federal Trade Commission announced the order on 19 December 2023. Its release states, "The system generated thousands of false-positive matches."

The cost of a wrong match lands on a person. According to the US Federal Trade Commission, customers were "erroneously accused by employees of wrongdoing" after false matches. The release also lists what was missing. First, Rite Aid did not test the system's accuracy before use. Second, it did not track the rate of false matches after rollout. Finally, it did not train staff to expect wrong matches.

Every one of those gaps is a process gap, not a camera gap. Because staff act on every alert they get, a false alert is never free. Therefore a loss camera needs a measured false-alert rate and a review step before anyone approaches a customer. Also, cameras that see faces carry privacy duties that a shelf camera never does, so privacy design comes before go-live.

Why does the integration cost more than the cameras?#

Recognition is a small line in a retail vision budget, priced per thousand units, so the catalogue, the store events and the staff tasks around it are the bill. Google's Cloud Vision pricing page, read on 1 October 2026, counts each feature applied to an image as a unit. It lists label detection as free for the first 1000 units a month. Above that, the price for each further block of 1000 units is $1.50.

Show data table
Recognition is priced in dollars per thousand units. Source: Google Cloud Vision pricing page, accessed 1 October 2026
Item Value
Units 1001 - 5,000,000 / month 1.5
Units 5,000,001 and higher / month 1

Past the free first 1000 units, label detection costs $1.50 per 1000 units, and $1.00 above 5,000,000 a month.

Figure Recognition is priced in dollars per thousand units. Source: Google Cloud Vision pricing page, accessed 1 October 2026 Google Cloud, accessed 1 October 2026

Instead, the money goes to four places. First comes the catalogue. Google's Vision API Product Search docs say retailers "create products, each containing reference images that visually describe the product from a set of viewpoints". As a result, every new product and pack design needs those images before a camera can name it.

Second come the store events. A recognised product means nothing until it becomes a POS line, a stock change or a label check. Third comes the task flow, the screen or handheld where staff see the alert and close it. In practice, that loop is integration work across POS and inventory, not model work.

Fourth comes the exit plan. Microsoft's migration guide says the Image Analysis API "will be retired on September 25, 2028". So a project built on one vendor's service needs a budget to move when that service ends.

Where the money goesA recognised product is worth nothing until it passes the catalogue, becomes a store event, reaches a staff task and survives the vendor's retirement date. Sources: Google Vision API Product Search docs; Microsoft migration guide (retirement 25 September 2028).

What should a pilot prove before a store chain scales it?#

A pilot should prove three numbers before scale: the false-alert rate staff will tolerate, the share of alerts that become finished tasks, and the recovered sales against the loss. A demo that names products on a shelf only proves the camera can see.

The first number is the false-alert rate, measured in the store over weeks. The Rite Aid order shows what happens when it goes untracked. So set the rate staff will accept before the pilot starts, then count every alert that turns out wrong.

The second number is alert-to-task completion. Because three in four gaps start in store practice, the alert is only worth something once a person restocks the shelf. So count how many alerts became closed tasks, and how long each took.

The third number is recovered sales. Compare pilot stores with similar stores that have no cameras, over the same weeks. Then set the gain against the store's own loss from the calculator above. For instance, say a store with 800,000 dollars of out-of-stock loss wins back a tenth of it. That example earns 80,000 dollars a year, which has to cover cameras, recognition, catalogue work and staff time.

  1. Number 1

    The false-alert rate

    Measured in the store over weeks, against a rate staff agreed to accept before the pilot started. The FTC order of 19 December 2023 shows what happens when it goes untracked.

  2. Number 2

    Alert-to-task completion

    How many alerts became closed tasks, and how long each took. Three in four gaps start in store practice (Gruen and Corsten, 2007), so an alert is worth something only once a person acts.

  3. Number 3

    Recovered sales

    Pilot stores against similar stores with no cameras over the same weeks, set against the store's own loss from the calculator above.

If any of the three numbers is missing, the pilot has not answered the question. In short, scale on the numbers, never on the demo.

When is a camera the wrong tool for the job?#

When a barcode, an RFID tag or a scan-as-you-go cart already answers the question, a camera adds cost and error without adding an answer. The rule is simple. If the item can identify itself, let it.

Amazon follows that rule in its own stores. For apparel, for instance, its Just Walk Out FAQ describes RFID lanes, where "the RFID tags in the clothing and other apparel is read by RFID readers" at the exit gate. For large grocery trips, it moved shoppers to Dash Cart, where the shopper scans and weighs items as they go. As a result, both choices take the hardest part of the job away from the camera.

So three cases point away from vision. First, if stock counts are the goal and items carry tags, an RFID read is more direct than a camera count. Second, if checkout is the goal and baskets are large, a scanning cart or a self-checkout lane is cheaper to make right. Third, if the store's stock records are wrong, fix receiving and counting first, because a camera will only report gaps the records already hide. In each case, the simpler tool wins.

Three cases where a simpler tool than a camera answers the question. Sources: Amazon Just Walk Out FAQ (RFID lanes) and Amazon's Dash Cart article, updated 27 January 2026.

GoalWhenThe simpler tool
Stock countsItems carry tagsAn RFID read
CheckoutBaskets are largeA scanning cart or a self-checkout lane
Accurate gap reportsThe stock records are wrongFix receiving and counting first

Where should a retail vision project go next?#

Start from the store system the detection must reach, POS, inventory or the task list, and choose the camera use case that system can act on. That order keeps the project tied to the money from the start.

When the detection must reach the till, read how cloud POS systems handle retail data before choosing a camera. When it must reach stock and reordering, how supply chain management software tracks inventory covers the records a shelf alert has to update. The same pick-the-one-that-pays-first method applies on a factory floor, as AI software in manufacturing shows.

Where a vendor's product stops at the alert, the missing piece is the store-side workflow, which is custom software built around the task list. Even so, the sources linked above, from the NRF survey to the RPC paper, are enough on their own to size computer vision in retail for one store and set a pilot's three numbers.

Questions this post answers

Which computer vision in retail uses actually pay back?
Computer vision in retail works where a missed detection costs little and feeds one task list, and stalls where every item must be right and must reach payment.
Why does checkout-free computer vision in retail struggle with large baskets?
On the RPC checkout dataset, the best method got the whole basket right 73.17 percent of the time with light clutter and 42.48 percent with heavy clutter. Checkout accuracy is the product of every item being right, so it falls as baskets grow.
What should a computer vision in retail pilot prove before a chain scales it?
A pilot should prove three numbers before scale: the false-alert rate staff will tolerate, the share of alerts that become finished tasks, and the recovered sales against the loss. If any of the three numbers is missing, the pilot has not answered the question.

Keep reading