When Big Data Enters the Kitchen: Letting Big Data Teach You to Cook!
A word before we start
Food (image from the internet)
In the author's self-introduction in the blog's sidebar, one line reads "kitchen enthusiast." Although I'm not much of a cook, the kitchen really is one of my hobbies. Of course, I have many hobbies—mathematics, physics, astronomy, computer science—I like and want to learn all of them, ending up broad but not deep. As mentioned in earlier posts, data mining is also one of my hobbies, and when the two hobbies of data mining and cooking meet, what interesting results might come of it?
That's exactly what I did: from the "Home-style Dishes" section of Meishi China (美食中国), I wrote a simple crawler to scrape a batch of recipe data, and then ran a simple analysis on it. (My sincere thanks to Meishi China here. I chose it because its data is fairly well-structured.) I did the analysis on my company's high-performance server, which made the whole process especially pleasant.
In total I collected 18,209 recipes, containing 9,700 distinct ingredients (including main ingredients, auxiliary ingredients, and seasonings—some may be duplicated due to inconsistent naming, etc.). Of course, compared to the "big data" standards of many fields, this amount of data is really nothing to brag about. But in the kitchen domain, which big data rarely touches, it should count as fairly substantial. more
Simple statistics
The simplest thing to do is a statistical analysis of the ingredients. Can you guess what appears most often?
You don't need to be a food expert to guess: the most frequent item is definitely salt! Salt is often called "the king of flavors," and very few dishes go without it. Next comes cooking wine, then light soy sauce—all seasonings and condiments. This also shows that Chinese cuisine places great emphasis on ingredients, with all sorts of condiments filling the list. Among main ingredients, potato ranks 28th, pork belly ranks 38th, and so on.
Salt 11200
Cooking wine 4601
Light soy sauce 4413
Ginger 3671
Scallion 2854
Chicken essence 2579
White sugar 2440
Sugar 2303
Oil 2297
Garlic 2058
Egg 1924
Soy sauce 1883
Dark soy sauce 1625
White pepper 1619
Sichuan pepper 1571
Carrot 1324
......
Word2Vec results
If we treat each recipe as a pre-segmented "sentence," we can use this "corpus" to train a Word2Vec model and see what interesting results come out. (Even if nothing interesting turns up, that's fine—this is exploratory work.) The entire training process was surprisingly fast, taking almost less than a second.
For readers unfamiliar with it: Word2Vec is a model that converts words into real-valued vectors, since a word must be converted into numbers before a computer can process it. The vectors produced by Word2Vec also have a special property: the cosine similarity between two word vectors represents how similar the two words are.
After training the Word2Vec model, the first thing we can do is compare the similarity between two words. Some results are quite ordinary, for instance:
pd.Series(model.most_similar(u'pork belly'))
0 (spare ribs, 0.882662177086)
1 (skin-on pork belly, 0.866969347)
2 (dried string beans, 0.864805340767)
3 (quail eggs, 0.850470840931)
4 (pickled cabbage, 0.842567443848)
5 (duck leg, 0.841659963131)
6 (three-yellow chicken, 0.837065219879)
7 (old brine broth, 0.828875720501)
8 (chicken gizzard, 0.827436089516)
9 (crucian carp, 0.826281666756)
But some results are quite surprising, for instance:
pd.Series(model.most_similar(u'chicken'))
0 (corn, 0.939546108246)
1 (tea tree mushroom, 0.914446234703)
2 (sweet corn, 0.888315618038)
3 (fresh shrimp, 0.88096922636)
4 (crab-flavor mushroom, 0.870144784451)
5 (red carrot, 0.86743336916)
6 (spaghetti, 0.864846467972)
7 (Kewpie salad dressing, 0.860477805138)
8 (pork loin, 0.85995388031)
9 (white mushroom, 0.855247914791)
Here, "chicken" and "corn" turn out to be highly similar! This suggests there must be an untold story between them.
Where does this come from? The Word2Vec model works based on word co-occurrence, so the reason for this phenomenon is likely: 1) corn and chicken are often cooked together; 2) corn and chicken are often cooked separately with similar ingredients. Indeed, on closer inspection, both turn out to be true—they are often used to make soup, and their accompanying ingredients are similar:
Recipes containing chicken (partial list):
140 [chicken, maca, goji berries, red dates, longan, lotus seeds, ginger]
144 [chicken, pointed pepper, vegetable oil, old ginger, garlic, dried chili, Sichuan pepper, salt, cooking wine, light soy sauce, white sugar]
267 [chicken, potato, green pepper, onion, flour, ginger and garlic, small chili, dark soy sauce, white sugar, light soy sauce, Sichuan pepper and star anise...
313 [chicken, beer, potato, Sichuan pepper, star anise, small chili, ginger and garlic, light soy sauce, dark soy sauce]
520 [chicken, breadcrumbs, egg, glutinous rice flour, light soy sauce, oyster sauce, salt, white pepper]
961 [cucumber, chicken, dried shrimp, red dates, ginger, star anise, scallion]
1005 [glutinous rice, chicken, shrimp, shredded ginger, chopped scallion, light soy sauce, oyster sauce, white pepper, starch]
1095 [chicken, egg white, flour, fresh ginger, garlic, cooking wine, salt, black pepper]
1178 [chicken, milk, salt, white pepper, garlic powder, low-gluten flour, starch, ice water, crushed peanuts, cooking oil, wheat...
1551 [chicken, king oyster mushroom, onion, dried chili, scallion, ginger, garlic, Sichuan pepper, rock sugar, salt]
Recipes containing corn (partial list):
106 [chicken wings, 2–3 pieces of spare ribs, salt, ginger, celery, cooking wine, carrot, corn, red dates, American ginseng slices, caterpillar fungus...
172 [Chinese yam, corn, spine bone, salt, ginger]
316 [pork bones, corn, iron-rod yam, red dates]
441 [corn, tomato, tofu, vegetable oil, old ginger, scallion, Sichuan pepper, salt, beef powder]
450 [Chinese yam, spare ribs, carrot, fresh ginger, cooking wine, thirteen-spice powder, goji berries, corn, scallion segments, star anise, refined salt]
483 [red carrot, pumpkin, celery, corn, broccoli, pine nuts]
485 [lotus root, carrot, shiitake mushroom, peanuts, red dates, corn, ginger slices]
509 [spare ribs, corn, king oyster mushroom, goji berries, bay leaf, MSG, salt]
789 [chicken wings, enoki mushroom, shiitake mushroom, white mushroom, brown beech mushroom, snow mushroom, scallops, caterpillar fungus, corn, salt, ginger...
828 [diced meat, corn, carrot, fish sauce, light soy sauce, salt, peanut oil]
It seems our little experiment did produce some genuinely interesting results. From experience alone, it's not easy to notice a correlation between "chicken" and "corn," yet through data mining, given enough data, such interesting patterns can indeed be discovered. There are similar results elsewhere: the similarity between beef and squid reaches 96%, and between beef and potato it's 91%! Here's the data—see if you can explain the beef–squid connection yourself:
Recipes containing beef (partial list):
46 [beef, carrot, chopped scallion, curry powder, salt, coconut milk, potato, onion, ginger, Korean soy sauce, Thai fish sauce]
70 [beef, wood ear mushroom, carrot, red bell pepper, green pepper, chopped scallion, shredded ginger, minced garlic, peanut oil, Pixian broad bean paste,...
148 [beef, green pepper, onion, ginger and garlic, Sichuan pepper, bay leaf, star anise, dark soy sauce, light soy sauce, cumin powder, white pepper]
272 [beef, taro, star anise, bay leaf, cassia bark, Sichuan pepper, ginger, chili, rock sugar, scallion, salt, cooking wine, light soy sauce]
290 [beef, carrot, onion, red wine, meat broth, salt, white pepper, tomato paste, butter, flour, bay leaf...
404 [beef, bay leaf, star anise, Sichuan pepper, ginger and garlic, oyster sauce, light soy sauce, salt, chicken essence]
433 [bean sprouts, beef, scallion, green pepper, oyster sauce, salt]
452 [beef, white granulated sugar, Sichuan pepper, sauce paste, salt, carrot, soy sauce, star anise, garlic, cooking wine]
455 [beef, dried yellow paste, thirteen-spice powder, cinnamon powder, star anise powder, Sichuan pepper powder, ginger powder, salt, old broth]
534 [potato, beef, salt, cooking wine, light soy sauce, garlic, fresh ginger, cilantro]
Recipes containing squid (partial list):
187 [squid, spare-rib sauce, sugar, oyster sauce, onion, chili, bell pepper]
284 [squid, bell pepper]
374 [squid, onion, white sesame, garlic chili paste, scallion, ginger, chopped scallion, barbecue sauce]
996 [onion, luffa, squid, oil, salt, soy sauce, white sugar, cooking wine, oyster sauce]
1468 [squid, onion, lettuce, premium soy sauce, garlic chili sauce, sugar, oyster sauce, cooking wine, salt, chicken essence]
1502 [squid, green and red chili, ginger and garlic, salt, white sugar, light soy sauce, starch, Sichuan pepper]
1577 [squid, bell pepper, soybean paste, ginger]
1619 [shrimp, white clam, squid, straw mushroom, tom yum paste, fresh lime leaf, fish sauce, coconut milk, sugar]
1796 [razor clam, squid, chives, white pepper, cooking wine, salt, ginger]
1798 [squid, cumin, salt, peanut oil]
Apriori association rules
Another potentially meaningful thing to try is mining association rules. Given the modest size of the data, I simply used the Apriori algorithm.
Before mining the rules, I first did some preprocessing: 1) removed salt, since its count is so large that if left in, most of the mined rules would include salt—which would essentially just be telling us "remember to add salt when cooking," a rather meaningless rule; 2) removed ingredients that appear only once, since these carry too little information and basically never show up in the rules—keeping them would only increase computation.
After this preprocessing, requiring a support (the frequency of the rule) of 0.01 and a confidence (the reliability of the rule) of 0.8, we get the following rules:
Rule Support Confidence
Cooking wine--Scallion--Garlic--Ginger 0.019935 0.912060
Sugar--Scallion--Garlic--Ginger 0.011203 0.879310
Cooking wine--Sichuan pepper--Scallion--Ginger 0.010544 0.872727
Cooking wine--Dark soy sauce--Scallion--Ginger 0.011643 0.868852
Star anise--Scallion--Ginger 0.016695 0.858757
Cooking wine--Sugar--Scallion--Ginger 0.013345 0.846690
Cooking wine--Light soy sauce--Garlic--Ginger 0.013345 0.840830
Light soy sauce--Scallion--Garlic--Ginger 0.012741 0.840580
Dark soy sauce--Garlic--Ginger 0.013290 0.831615
Sichuan pepper--Scallion--Ginger 0.019221 0.825472
Cooking wine--Garlic--Ginger 0.032511 0.821082
Cooking wine--Light soy sauce--Scallion--Ginger 0.019057 0.816471
Dark soy sauce--Scallion--Ginger 0.017903 0.808933
Cooking wine--Scallion--Ginger 0.050799 0.805749
What these mean is: if cooking wine, scallion, and garlic show up, then ginger should be added too; if sugar, scallion, and garlic show up, remember to add ginger too; and so on. All these rules end with ginger, telling us when ginger is needed—rules that are pretty useful in cooking (for beginners, at least). This also shows that ginger is a very important ingredient in Chinese cooking.
We can loosen the conditions a bit to try to mine more rules. Lowering the confidence threshold to 0.7, we get:
Rule Support Confidence
Scallion--Garlic--Ginger 0.038497 0.799316
Cassia bark--Bay leaf--Star anise 0.018013 0.782816
Rock sugar--Cassia bark--Star anise 0.010160 0.780591
Sichuan pepper--Garlic--Ginger 0.014169 0.779456
Light soy sauce--Bay leaf--Star anise 0.010874 0.770428
Cooking wine--Cassia bark--Star anise 0.014938 0.764045
Cassia bark--Star anise 0.031633 0.761905
Cassia bark--Sichuan pepper--Star anise 0.015267 0.761644
Dark soy sauce--Bay leaf--Star anise 0.011423 0.759124
White pepper--Scallion--Ginger 0.014498 0.758621
Cassia bark--Dark soy sauce--Star anise 0.013070 0.757962
Sugar--Scallion--Ginger 0.022297 0.753247
Cassia bark--Light soy sauce--Star anise 0.011148 0.751852
Cooking wine--Bay leaf--Star anise 0.012411 0.750831
Sichuan pepper--Bay leaf--Star anise 0.014608 0.745098
Scallion--Vinegar--Ginger 0.011203 0.744526
Rock sugar--Light soy sauce--Dark soy sauce 0.010929 0.742537
Bay leaf--Star anise 0.028777 0.738028
Ginger--Cassia bark--Star anise 0.011917 0.735593
Starch--Scallion--Ginger 0.011038 0.730909
Ginger--Bay leaf--Star anise 0.010215 0.723735
Ginger--Sugar--Garlic--Scallion 0.011203 0.720848
Sugar--Garlic--Ginger 0.015542 0.712846
Light soy sauce--Scallion--Ginger 0.032402 0.704898
Scallion--Soy sauce--Ginger 0.017244 0.700893
If the rules about ginger still seem too ordinary, then these rules should feel more meaningful. For example, "cassia bark--bay leaf--star anise," "rock sugar--cassia bark--star anise," "cassia bark--light soy sauce--star anise," and so on—these combinations should correspond to recipes for braised ("lu wei") dishes. Not every home cook knows these combinations off the top of their head, but through association rule mining, they can be uncovered.
There are also some rules with even higher confidence but slightly lower support:
Rule Support Confidence
Cooking wine--Dark soy sauce--Scallion--Garlic--Ginger 0.005602 0.953271
Cooking wine--Sichuan pepper--Scallion--Garlic--Ginger 0.005547 0.952830
Cooking wine--Sugar--Scallion--Garlic--Ginger 0.007634 0.952055
Cooking wine--Cassia bark--Scallion--Ginger 0.005711 0.936937
Cooking wine--Light soy sauce--Scallion--Garlic--Ginger 0.007743 0.921569
Cassia bark--Sichuan pepper--Scallion--Ginger 0.005217 0.913462
Cooking wine--White sugar--Garlic--Ginger 0.005931 0.805970
Cooking wine--Scallion--Ginger 0.050799 0.805749
Rock sugar--Dark soy sauce--Bay leaf--Star anise 0.005162 0.803419
These are more detailed and precise seasoning formulas. Note that these are results automatically mined by the computer—the computer is our master chef here.
In closing
This post tries to bring together two of my own interests—data mining and cooking—and see what interesting results emerge. Indeed, some seemingly interesting results did come out; of course, these also reflect some of my own understanding of cooking, but whether they're truly interesting is for the reader to judge.
This kind of mining is essentially text mining, or could be classified under natural language processing. As this post shows, the methods used are basically NLP methods. Readers familiar with the field will know that the difficulty in NLP lies in feature construction—that is, how to represent a word numerically. This post is one attempt at that, but due to insufficient data and other factors, the conclusions aren't necessarily accurate. This attempt isn't a resounding success, but the process is worth referencing. Perhaps with more data, the value of the findings could grow.
Intuitively, this kind of mining is meaningful—we can uncover information we didn't know from an ordinary domain, even if it's information we technically already "knew" but simply never noticed. The computer helps us discover it, allowing us to properly recognize it, or make better use of it. Data mining can help us live better. Indeed, data mining technology ought to be democratized and made accessible to everyone, because our own daily lives are our most important source of data.
Finally, here is the scraped data for anyone interested: recipe_data.zip
Translated automatically with claude-sonnet-5; all equations are reproduced verbatim from the source. Copyright remains with the original author.