DH 2026

Daejeon, July 27–31

Fri, July 3109:00–10:30S104204-205
Short Paper

“Add Y% tbs. chopped pints of chicken stock.” A small data, low-tech approach to the computational analysis of labour in historical cooking recipes

Ginevra Sanvitale
Trinity College Dublin, Ireland · sanvitag@tcd.ie

Introduction. From Food Computing to Digital Food History

Computational analysis of cooking recipes largely stem from Food Computing, but the field’s methods and results can be challenging to integrate into other disciplines, especially historiography. Food Computing typically relies on digitally born datasets of contemporary recipes (Gilal et al. 2024), which are of limited interest to historians. The field mostly focuses on developing, testing and improving data retrieval and processing methods (Wu et al. 2025; Bagler and Goel 2024; Min et al. 2020), informed by computational questions rather than conceptual or analytic. My on-going project starts instead from a historiographic question—How did the industrialisation of poultry farming change labour requirements for chicken-based recipes across time?—that includes a computational one: How to calculate the human labour required to perform a cooking recipe?

Few studies attempted to model the labour intensity (or rather, difficulty) of recipes. Yajima and Kobayashi (2009) introduced a system based on seasoning, ingredients, and cooking methods, using frequency-based scores and manually classifying cooking verbs as “easy” or “laborious.” Kusu et al. (2017) refined it by defining four verb-difficulty levels from Japanese home-economics textbooks. They also added a quantitative validation method, through a questionnaire in which 140 participants ranked recipes’ difficulty. A qualitative validation example comes from Fukumoto et al. (2020), who assessed recipe instructions’ effectiveness by having eight cooks prepare dishes, record preparation times, and discuss their experiences. However, these studies offer cuisine-specific, contemporary-oriented, and technically complex methodologies that are difficult to integrate across disciplines and cultural contexts.

I test and validate a more extended, versatile linguistic model, reframing this work from “Food Computing” to what I term “Digital Food History.” This notion builds on Digital Food Studies (Leer and Krogager 2021),research that is in conversation with the social sciences and humanities rather than computer science. History-driven research with a quantitative focus or advanced computational workflows can fall under Digital Food History too (Blankenship 2024, 2025; Ciambella 2023; Westbrook et al. 2023). But shifting the focus from “computing” to “history” also aims to foster the inclusion of small data, low-tech approaches which are accessible to a wider pool of humanities scholars.

I work on two case studies. The first is based on Italian-American cookbooks in English, where I analyse changes to labour intensity in poultry-based recipes, and compare them with non-poultry recipes: Simple Italian cookery (1912, 119 recipes), Economical Italian cookbook (1917, 67 recipes), Practical Italian recipes for American kitchens (1918, 50 recipes), The Italian cook book (1919, 180 recipes), The European cookbook for American homes (1936, 87 recipes), Specialitá culinarie italiane (1936, 119 recipes); The art of Italian cooking (1948, 196 recipes); The Talisman of Happiness (1950, 650 recipes); The new Italian cookbook (1959, 178 recipes). The second case study is based on Italian language poultry-based recipes from cooking magazineLa Cucina Italiana (1929-1975, 543 recipes), allowing to track variations within a single, coherent corpus.

A small data, low-tech approach to measure recipes’ labour intensity

My labour intensity model incorporates tools and action-qualifying adverbs alongside ingredients and verbs. Tools matter because they must be cleaned, adding labour, while adverbs specify how an action is performed, increasing effort. I avoid verb-difficulty categories, as they vary by ingredient and cuisine. I measure labour intensity by summing: 1) ingredients, weighted by their corpus frequency; 2) tools; 3) cooking actions; and 4) adverbs, counted at half-weight since they modify an action rather than add a new one. Key challenges requiring additional manual curation include scoring ingredient substitutions (“oil or butter”), disambiguating terms (“brown” as verb or adjective), detecting cross-recipe references (“prepare as in recipe X”), and inferring implicit tools (e.g., instructions to “chop” without mentioning a knife).

I extract items using a wordlist-based approach. Although Food Computing often relies on Named Entity Recognition for ingredient retrieval (Agarwal et al. 2024), this is not ideal for historical recipes because of OCR noise (which often appears around ingredients) and datasets training biases. Conversely, wordlists work well for thematic corpora with limited lexical variation and can cover all target categories—ingredients, tools, verbs, and adverbs. I initially explored datasets CulinaryDB (Bagler 2017), FoodOntoMap (Popovski et al. 2019), EPIC-KITCHENS (Damen et al. 2021) and CookIT (Artese et al. 2019). But these were either too narrow or too broad, requiring extensive additional processing. I therefore built new wordlists using a corpus-driven method. For English, I used EPIC-KITCHENS’ cooking-actions dataset as a base for verbs and De Sola and De Sola’s (1969) historical cooking vocabulary for ingredients and tools. I then queried the corpora with LancsBox’s Words tool (with SpaCy medium model for POS tagging and lemmatisation) to identify candidate verbs, adverbs, and noun–adjective pairs signalling ingredients or tools, and manually curated the results. The final lists contain 454 ingredients, 81 tools, 300 verbs, and 73 adverbs. For Italian, I follow the same procedure but using TreeTagger, which outperforms SpaCy for Italian POS tagging and lemmatisation (Artese and Gagliardi 2019).

For the final processing, I use two Python scripts with the csv, regex, and Counter modules to: 1) find ingredients and tools (counted once per recipe), verbs and adverbs (counted per occurrence); 2) compute ingredient frequencies across the corpus. The output is saved as a CSV file, where I calculate recipes’ labour-intensity score using the formula above. Following Kusu et al. (2017) and Fukumoto et al. (2020), I qualitatively validate the model by cooking recipes of varying difficulty, noting preparation time and task complexity, and comparing these observations with the computational output. I then interview professional Italian cooks to rank the recipes by difficulty and comment on the model and my cooking experience.

Conclusion

Preliminary analysis of the English language corpora suggests that verbs are key indicators of labour intensity, while adverbs mostly reflect authors’ stylistic choices. Chicken recipes increased over time and were generally of medium to high difficulty. Final analysis will assess whether the industrialisation of poultry farming reduced family cooking labour or, as in Schwartz Cowan (1983), resulted in “more work for mother.”

References
  1. Agarwal, Ayush / Kapuriya, Janak / Agrawal, Shubham et al. (2024): “Deep Learning Based Named Entity Recognition Models for Recipes”, in: Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024): 4542–54.
  2. Artese, Maria Teresa / Ciocca, Gianluigi / Gagliardi, Isabella (2019): “Cookit: A Web Portal for the Preservation and Dissemination of Traditional Italian Recipes”, in: International Journal of Business, Human and Social Sciences 13, 2: 171–76.
  3. Artese, Maria Teresa / Gagliardi, Isabella (2019): “Preprocessing Pipeline for Italian Cultural Heritage Multimedia Datasets”, in: Archiving Conference 16, 1: 81–85 https://doi.org/10.2352/issn.2168-3204.2019.1.0.18.
  4. Bagler, Ganesh (2017): “CulinaryDB” https://cosylab.iiitd.edu.in/culinarydb/ [06.05.2026].
  5. Bagler, Ganesh / Goel, Mansi (2024): “Computational Gastronomy: Capturing Culinary Creativity by Making Food Computable”, in: NPJ Systems Biology and Applications 10, 1: 72 https://doi.org/10.1038/s41540-024-00399-5.
  6. Blankenship, Avery (2024): “What We Didn’t Know a Recipe Could Be: Political Commentary, Machine Learning Models, and the Fluidity of Form in Nineteenth-Century Newspaper Recipes”, in: Journal of Cultural Analytics 9, 1 https://doi.org/10.22148/001c.115371.
  7. Blankenship, Avery (2025): “Literary Forensics as Method: Chemical Analysis, Food Stains, and Readerly Encounters with Nineteenth-Century Cookbooks”, in: American Literature 97, 3: 409–38 https://doi.org/10.1215/00029831-11934312.
  8. Ciambella, Fabio (2023): “A Corpus-Driven Analysis of the Early Modern English Recipe Manuscripts at the Folger Shakespeare Library: Zooming in on Morphosyntax and Pragmatic Interfaces”, in: Status Quaestionis 25. https://doi.org/10.13133/2239-1983/18570.
  9. Damen, Dima / Doughty, Hazel / Farinella, Giovanni Maria et al. (2021): “The EPIC-KITCHENS Dataset: Collection, Challenges and Baselines”, in: IEEE Transactions on Pattern Analysis and Machine Intelligence 43, 11: 4125–41 https://doi.org/10.1109/TPAMI.2020.2991965.
  10. De Sola, Ralph / De Sola, Dorothy (1969): A Dictionary of Cooking. Meredith Press.
  11. Fukumoto, Hayate / Ohsugi, Takafumi / Matsushita, Mitsunori (2020): “Presenting Action-Centered Recipe to Reduce Cooking Failure for Beginners”, in: The 34th Annual Conference of the Japanese Society for Artificial Intelligence.
  12. Gilal, Nauman Ullah / Al-Thelaya, Khaled / Al-Saeed, Jumana Khalid et al. (2024): “Evaluating Machine Learning Technologies for Food Computing from a Data Set Perspective”, in: Multimedia Tools and Applications 83, 11: 32041–68 https://doi.org/10.1007/s11042-023-16513-4\.
  13. Kusu, Kazuma / Makino, Nozomi / Shioi, Takamitsu / Hatano, Kenji (2017): “Calculating Cooking Recipe’s Difficulty Based on Cooking Activities”, in: Proceedings of the 9th Workshop on Multimedia for Cooking and Eating Activities in Conjunction with The 2017 International Joint Conference on Artificial Intelligence, 20 August: 19–24 https://doi.org/10.1145/3106668.3106673.
  14. Leer, Jonatan / Krogager, Stinne Gunder Strøm (2021): Research Methods in Digital Food Studies. Routledge https://doi.org/10.4324/9781003010845.
  15. Min, Weiqing / Jiang, Shuqiang / Liu, Linhu / Rui, Yong / Jain, Ramesh (2020): “A Survey on Food Computing”, in: ACM Computing Surveys 52, 5: 1–36 https://doi.org/10.1145/3329168.
  16. Popovski, Gorjan / Koroušić Seljak, Barbara / Eftimov, Tome (2019): “FoodOntoMap Version 2: Linking Food Concepts across Different Food Ontologies”. In: Proceedings of the 11th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management (IC3K 2019): 195–202 https://doi.org/10.5281/zenodo.3600619.
  17. Schwartz Cowan, Ruth (1983): More work for mother: The ironies of household technology from the open hearth to the microwave. Basic Books.
  18. Westbrook, John D. / Jakacki, Diane / Heintzelman, Rebecca / Harnood, Juliya (2023): “Digitizing Suzette: Creating a Framework for the Collaborative Analysis of an Historical French Textbook”, in: Digital Humanities Conference, Graz, Austria, 2023.
  19. Wu, Yao / Kang, Ling / Guo, Quan / Zhao, Lei (2025): “A Review of Recipe Recommendation Methods”, in: IEEE Access 13: 92483–94 https://doi.org/10.1109/ACCESS.2025.3572149.
  20. Yajima, Asami / Kobayashi, Ichiro (2009): “‘Easy’ Cooking Recipe Recommendation Considering User’s Conditions”, in: 2009 IEEE/WIC/ACM International Joint Conference on Web Intelligence and Intelligent Agent Technology: 13–16 https://doi.org/10.1109/WI-IAT.2009.219.