joel taylor pedrós
blog

supermercapy: compare supermarket prices at 14 chains with python

diagram of one call to search_all fanning out into fourteen lines, one per supermarket, each labelled with the technology behind its site.

plusfresc and condis had the same offer on the 33 cl can of estrella damm, the second one at half price. plusfresc listed it at €0.67, and condis at €0.89.

plusfresc gives the average of the two cans as the price, (0.89 + 0.445) / 2 = 0.6675, which it shows as 0.67, with 0.89 as the previous price. condis gives the price of a single can and writes the offer next to it, "segunda unidad 50%". if you buy two, they come out the same at both shops. if you want one, the cheapest was €0.75, at carrefour, eroski, caprabo and bonpreu.

these are prices from 6 october for postcode 08013, and they come from supermercapy, a python library to compare supermarket prices: it reads the online shops of 14 chains.

from mercapy to supermercapy

in may 2024 i published mercapy, a client for mercadona's online shop. in september 2026 i rewrote the whole thing, and it's now the base of the program that reads mercadona's prices in 174 zones every day. supermercapy does the same with 14 chains: mercadona, consum, plusfresc, bonàrea, carrefour, lidl, bonpreu, dia, eroski, caprabo, aldi, ahorramás, alcampo and condis. the mercadona client is a port of mercapy. you install it with pip install supermercapy, and the code is at github.com/jtayped/supermercapy.

each chain has a class that inherits from a base client. searching, reading a product and reading the categories work on all of them. the full catalogue, offers, barcodes or nutritional information depend on the shop, and each client declares what it has. if you ask for something a shop doesn't publish, the library fails before making a single request.

how the websites of 14 supermarkets are read

none of these shops has a documented api, and each one does its own thing:

shopwhere the data comes fromwhat the price depends on
mercadonaa json api and an algolia indexthe warehouse
aldian algolia index per regionmainland spain, the balearics or the canaries
carrefouran empathy index and the page statethe point of sale
condisan empathy index and next.js pagesthe centre changes the range, not the price
bonpreuocado smart platformnothing, there's only one region
alcampoocado smart platformthe region
eroski, capraboapache tapestry htmlthe shop set by the anonymous session
ahorramássalesforce commerce cloud htmlnothing
bonàreaform requestsnothing
plusfresca json api with a guest tokenthe fulfilment centre changes the range
consuma json api with two headersthe zone changes the range, not the price
lidlfour apisthe region
diaa json api with a session cookiethe postcode, almost never

at dia, of 871 products sold in both madrid and barcelona, only one had a different price.

the worst case is condis. to find out which centre serves a postcode, the site calls a next.js server action, and that action's id changes with every deploy and only appears in the page's scripts. the client reads the home page, going through an anonymous login, then the scripts in order until one names the action. in october that was 12 scripts and about 1.4 mb. in total, 18 requests to resolve one postcode. that's why the docs recommend resolving it once and saving the centre.

bonpreu and alcampo sit behind an aws firewall with a budget per ip address. at bonpreu, in october, it was 8 to 13 requests from scratch; at alcampo, 7 or 8 every half hour on the product service. after that they answer 202 with an empty body for tens of minutes. these two clients wait 2 seconds between requests and are meant for one-off lookups, not for crawling the catalogue.

at eroski, every product in the list carries an analytics attribute with the price as a float, which drops trailing zeros. the client reads the printed price and only uses the analytics one if there isn't one.

one price model

every client translates all of this into the same classes, immutable and using Decimal, never float. each product has the amount, the previous price, the reference price and the promotions.

the hard part is the reference price. mercadona sends units as "L" and "dc", consum as "1 L" and "100 Gr", bonpreu as "PER_1KG". aldi sends g and 100-g on prices that are actually per kilo: €1.59 for 200 g, with a reference price of 7.95. price.reference restates all of it per kilo, litre, piece, dose, metre or square metre, and None means unknown, not zero.

one call, fourteen threads

search_all() runs the same search on every shop, each in its own thread with its own client. given a postcode, the shops that need one resolve it before searching. the result always comes back in the same order, not the order they finished in, and each shop says how it went: it answered, it wasn't asked, it doesn't serve that postcode, or it failed. one shop failing doesn't hide what the others found.

from supermercapy import find_same, search_all

result = search_all("estrella damm", postal_code="08013", page_size=10)
can = next((s, p) for s, p in result.products
           if s == "carrefour" and "lata 33 cl" in p.name)
same = find_same(can, result.products, threshold=0.75)

the search for the can took 45 requests and 9 seconds, 19 of them for condis, and all 14 shops answered.

python code using supermercapy that calls search_all and find_same, and the terminal output: the 33 cl can of estrella damm at ten chains. plusfresc 0.67, was 0.89, second one at half price. carrefour, eroski, caprabo and bonpreu 0.75. alcampo 0.86. mercadona, consum, dia and condis 0.89, condis with segunda unidad 50 %. 14 of 14 shops answered.
run on 6 october 2026 at 19:04, postcode 08013.

finding the same product at every supermarket

only mercadona, consum, carrefour and lidl publish barcodes, and mercadona only on each product's page. to tell whether a can from carrefour is the same as one from bonpreu, you almost always have to look at the name, the brand and the size.

if both barcodes match, the score is 1. if not, it's half the similarity of the names, plus a quarter if the brand matches and a quarter if the size matches. a conflicting brand, size or packaging halves it. names are compared without the brand, the size or the packaging words, in spanish and in catalan, and only the differences count: a descriptor like "refresco" costs 0.15, any other word 0.7, and a number or a negation like "sin lactosa", 1. the similarity is 1 / (1 + cost).

with the brand and size right, a single word of difference leaves the similarity at 0.59 and the score at 0.79, just under the default threshold, 0.8. bonpreu and condis scored exactly that: brand and size right, name at 0.59, total 0.79. that's why i used 0.75. and i started from the carrefour can, which has a barcode, because starting from eroski's, at 0.75, consum's lemon beer slipped in. this way find_same() finds the can even though consum calls it "cerveza lata" and bonpreu "cervesa especial en llauna".

two line charts showing the precision and recall of score_same by threshold. on the 10 development searches, at a threshold of 0.75 precision is 0.75 and at 0.80 it's 1.00, while recall drops from 0.81 to 0.59 between 0.75 and 0.90. on the 4 held-out searches, precision goes from 0.76 to 0.996 and recall from 0.81 to 0.64.
0.8 is right on the edge: at 0.75 a one-word difference already gets through.

to choose the threshold i labelled 1,343 results by hand from 14 real searches on 4 october. at 0.8, score_same() never gets it wrong on the 10 searches i wrote the rules with, and it finds 73% of the pairs. on the 4 i labelled before running it, the first pass gave a precision of 0.96 and a recall of 0.55. for a category the rules have never seen, that's the honest estimate.

in the search for the can, it turned up at 10 of the 14 shops. lidl returned products that had nothing to do with it, like lamps and a christmas star, aldi returned nothing, ahorramás returned estrella galicia and mahou, and bonàrea only sells it in packs of 6 and 12.

prices ahead of time

lidl and aldi publish prices before they take effect. on sunday 4 october, 113 aldi products on the mainland already carried monday's price, and 74 carried a second one, for when the first ran out. at lidl, get_offers(week="next") reads the following week's offers days before they start.

when a site changes

each client has live tests that call every public method and compare the json of each response against a copy saved in the repository, so that a renamed field fails a test instead of silently turning into None. they run every week and open an issue for each shop that fails.

the audit on 4 october, before the first release, already found changes on several sites. lidl's search answered 406 to anyone asking for application/json, bonpreu's firewall blocked any client that didn't look like a browser, and mercadona had stopped sending the category level. carrefour lost a whole feature. its only complete listing was the sitemaps, and since october they answer with cloudflare's verification page. that's why supermercapy can read the full catalogue of 11 chains, not 14.