Daejeon, July 27–31
The use of metaphor in architecture has a long history [1], and metaphors are drawn from many source domains, including language, music, science, mechanics, and biology. While scholars disagree on whether to include or exclude highly normalized metaphorical terms adopted into professional language [2], it is widely known that metaphors comparing the built environment to the human body exist in the earliest written works in the western architectural discourse and that metaphors remain ubiquitous.
Biomimetic design involves the application of creative analogies from biology to the built environment and is typically driven by design intents related to sustainability and innovation [3]. There is much interest in biomimetic design, evidenced in the professional community by trainings on the topic [4], the use of metaphors in building rating systems [5], and evidenced in the scholarly community by the growing number of papers on biomimetic design [6]. Recent professional and scholarly attention given to biomimetic concepts has focused primarily on the deployment of these ideas to achieve novel technologies or higher performance [7,8,9].
What has received little systematic attention is an attempt to map these comparisons to human body systems and to analyze trends in their evolution over time. This study represents the initial steps to develop a workflow for collecting and organizing an architectural canon of writings and to find and extract these biological comparisons to identify trends in how they present and evolve.
To find and study body-to-built environment (BTBE) comparisons, we first assemble a corpus of written works in architecture. We use the final version (2000) of historian Charles Jencks’ landmark visualization [10, 11] of key people and movements, “Evolutionary Tree of 20th-century Architecture” (TET) first published in 1970 [12], as the basis for identifying works representing 20th-century architectural discourse.
Fig 1. Jencks’s “The Evolutionary Tree” as updated in 2000
From TET, we manually compiled 320 authors and their date placements and then cross-referenced the author list and date ranges against the HathiTrust Digital Library (HTDL) catalog containing over 18.6 million digitized volumes at the time of our search. We conduct fuzzy author-title matching within a date-limited subset of the full HTDL catalog, accessed via the HathiFiles TSV metadata dumps [13], using Levenshtein-based fuzzy string matching iteratively implemented in Python with the rapidfuzz library [14]. This task was difficult, time-intensive, and yielded many false positives, necessitating manual review of titles returned, finally yielding a list of 466 de-duplicated and relevant titles for our final digitized corpus.
We next curate a list of biological keywords, sourced from the National Library of Medicine’s Medical Subject Headings (MeSH) controlled vocabulary [15], Anatomy and Physiology for Health Professionals by Moini [16], and the influential article by Edward Trifonov, Vocabulary of Definitions of Life Suggests a Definition [17], establishing a list of 363 biological keywords. Our workflow begins with a full text search of each page for a biological keyword within a sentence and extracts this sentence and before and after sentences for context. We then search extracted passages for instances of built environment keywords, sourced from a manually compiled list of 72 relevant terms, to sort the results into high- and low-priority items for review (Figure 2). Our final KWIC corpus totaled 115,654 keyword sentences and associated context and metadata.
Fig 2. Workflow for identifying and extracting keywords-in-context from architectural volumes authored by individuals named in Jencks’s diagram.
The scale of keywords-in-context (KWIC) results made wholesale manual review unfeasible. Though we benchmarked automated review approaches, including using large language models (LLMs) to identify comparisons, such methods did not yield quality results. Instead, we focused on a case study subset of the keyword “skin” (with 1,178 KWIC results) and related terms “membrane,” “integument,” and “epidermis” (with a total of 349 results), which allowed us to manually verify each KWIC, leaving 1,017 sentences containing a keyword used in a verified BTBE comparison. Table 1 shows some example valid and invalid KWIC results.
| Keyword Sentence | Title | Publication Date | Author |
| True Positives | |||
| Each consuming monad would by the skin of its shelter, so to speak, capture, filter, transform, store, and consume that quantum of ENERGY needed by it, perhaps releasing some for collective use. | The sketchbooks of Paolo Soleri. | 1971 | Soleri, Paolo |
| The "modern" buildings of the period are distinguished by a few characteristic properties: they are usually derived from simple stereometric shapes; they appear as unitary volumes wrapped up in a thin, weightless SKIN of GLASS and plaster; and they show a puritan lack of MATERIAL texture and articulating detail. | Meaning in Western architecture | 1975 | Norberg-Schulz, Christian |
| False Positives | |||
| Hardly anyone swam in the ocean any more, as the water was full of jelly-fish that stung the SKIN badly. | Houses and people of Japan | 1937 | Taut, Bruno |
| Curious tan-gold foothills rise from the tattooed sand-stretches to join slopes spotted as the leopard-SKIN, with grease-bush. | Writings and buildings | 1960 | Wright, Frank Lloyd |
| It doesn't have the 'rubber nipples', the new nice-feel keys of this computer; but it is vaguely sensual, especially around the auditorium which, in SKIN tones, slithers and undulates its way to the ground. | Late-modern architecture and other essays | 1980 | Jencks, Charles. |
Table 1. Example positive and negative KWIC results in extracted sentences containing “skin,” “membrane,” “integument,” and “epidermis.”
Though this project is ongoing, we find that the count of skin-focused BTBE comparisons generally increases over the course of the 20th century before tapering off in the 21st century (Figure 3 illustrates full counts). We interpret this decline as largely due to the decrease of volumes published after 2000 in our dataset. Figure 4 shows the number of volumes, keyword sentences, and tokens per year in our corpus, illustrating the sparsity of data post-2000 and pre-1940 (which also surfaces in our “skin” BTBE data).
Fig 3. Frequency of “skin,” “membrane,” “integument,” “epidermis” in keyword sentences by year
Fig 4. Volume, KWIC, and token counts by year and log-scaled
The observed increase in “skin” BTBE comparisons raises the question of if such comparisons increasingly permeate the discourse in architecture–i.e. is the built environment described as being more dynamic or alive as time progresses? To test this, we defined sets of adjectives and verbs that denote dynamism/liveliness and stasis/inanimacy. We then generated word-level (word2vec) embeddings, implemented via Gensim [18], for each token within KWIC and classified each verb and adjective as either “dynamic” or “static” based on their distance between centroids of each class’s exemplar (anchor) words in vector space. We observe a general trend toward more static language post-1970, but then a narrowing of the gap, culminating in a shift back toward more balanced language as the 20th century progressed (see Figure 5). This signals an increased prevalence of biological description of the built environment in post-1990s architecture literature, a finding in-line with the documented increase of interest in biomimicry.
Fig 5. Net dynamism (count of dynamic minus count of static) of verbs and adjectives, by year, with 5-year rolling mean LOWESS trend line
| Static | Dynamic | |
| Verbs | cover, support, enclose, contain, stagnate, decline | change, evolve, respond, maintain |
| Adjectives | simple, dead, frozen, static, stable, nonbiological, inanimate, incapable | alive, biological, complex, capable, functional, variable, responsive, sensitive, bioclimatic, adjustable |
| Other | dying, stasis, constancy | living, process, metabolism |
Table 2. Anchor words used to classify KWIC verbs and adjectives as static or dynamic via distance in embedding space.
We find these results encouraging, though they are limited by the scale of the data–only 181 volumes are reflected in a subset of KWIC sentences comprising about 1% of our total KWIC data. As a result, for this work-in-progress, we are extrapolating larger trends in our corpus from a relatively small case study of a subset. Focusing on skin and related terms may also pose a unique linguistic challenge to our static versus dynamic exploration, as “skin” may appear as a verb or noun, a complexity that may not exist for other single-part-of-speech keywords. Further refinement of our data, including expanding this analysis to other terms and wider textual context, could better test this trend.
The opportunity to expand this work to more terms and more verified data also may uncover further interesting trends. We also plan to explore leveraging AI and LLMs for review tasks, as well as potentially engage with researchers active in architecture history and study for crowd-sourced annotation of results. Both our Jencks canon and our KWIC data are also well-positioned for further inquiry around any number of architectural topics, like sustainability, technological innovation, social concerns, or metaphors drawn from other source domains.
Finally, though our intent is to study one version of an architectural canon, the Jencks set is limited in incorporating important, diverse voices in the field. Roberts and Aiken created a similar chronogram that visualizes feminist spatial practices, including key figures, scholars, and movements [19]. We have run this diagram through the same process as TET and have extracted a small, verified corpus of works in the HTDL, and are poised to extend our investigation to this workset of volumes. The Roberts-Aiken corpus can serve as both a valuable corpus in which to investigate new research questions and an interesting comparison point to better verify the results and trends we observe in TET.