DH 2026

Daejeon, July 27–31

Poster

Poetry Annotator: A Structure-Preserving Web-based Tool Online-Offline Annotation of Underrepresented Poetry Collections

You Peng
University of Illinois Urbana-Champaign · youpeng2@illinois.edu
Gyuri Kang
Indiana University Bloomington · gyukang@iu.edu
William Sideri
University of Illinois Urbana-Champaign · wa7@illinois.edu
Kahyun Choi
University of Illinois Urbana-Champaign · kahyun@illinois.edu

Introduction

Carefully annotated datasets are essential for uncovering algorithmic bias and promoting debiasing techniques, especially in underrepresented domains that large-scale datasets tend to overlook (Yogarajan et al. 2023; Luka & Leurs 2023). Even at a small scale, structured annotated datasets allow computational analysis to “complement qualitative human readings” (Mauro 2014) and can help “pave the way for new annotations and new experiments” (Sprugnoli et al. 2023). Aligning with DH2026’s sub-theme "annotating | beyond patterns: interpretation with small data," this abstract introduces Poetry Annotator, an annotation tool developed within a broader research effort on underrepresented poetry. The tool supports building theme and emotion metadata of poems, including those with under-resourced languages (e.g., Native American and Hawaiian languages), to improve their visibility and support more equitable discovery in digital libraries. As an open-source, lightweight, and customizable tool, it can be easily adapted by other researchers for annotating their collections with different metadata schemes.

Background

We are annotating two poetry datasets to build richer metadata for underrepresented poetry and examine algorithmic biases in digital libraries. The poems were written by poets from five cultural groups, including Native American, Native Hawaiian and Pacific Islander, African American, Asian American, and Latino/a/x American, drawn from Poets.org and OCRed, copyright-protected poems accessed within the HathiTrust Research Center (HTRC) Data Capsule, a secure offline Linux environment. Dataset construction details are described in prior work (Choi & Kang 2025).

Existing annotation tools and methods were not compatible with our poetry annotation study. Simple text and spreadsheet tools flatten structure, which is problematic for poetry where lineation, spacing, and typography are essential to meaning-making (Bhyravajjula et al. 2025). They also lack administrative support for multi-annotator coordination. While annotators can read poems directly on the website (e.g., Poets.org), surrounding interface elements (e.g., navigation menus and advertisements) interrupt close reading and focused annotation. Furthermore, existing general-purpose annotation tools, such as BRAT, Doccano, TextAnnotator (Pontus et al. 2012; Hiroki et al. 2018; Abrami 2020; Colucci et al. 2024), likewise do not preserve poetic formatting or operate in secure offline environments, such as HTRC Data Capsules. We therefore developed a new annotation tool that preserves formatting, supports multiple annotators, and operates in a secure offline environment. After a two-month trial with 13 annotators and approximately 200 poems, the tool was refined and stabilized and is publicly available at: https://github.com/FAcc06/AnnotationTool_DH.

Key Features

Multidimensional annotation

Figure 1 shows the annotation interface design, where an annotator can assign theme and emotion tags to record their thematic conceptions of a poem and the emotions they feel while reading it. The interface shows 50 preset themes derived from Poets.org's top themes, offering annotators clear guidance and support for more consistent theme selection. It also includes eight preset emotions (anger, fear, anticipation, trust, surprise, sadness, joy, and disgust) from The NRC Emotion Lexicon (Mohammad & Turney 2013), a widely adopted emotion framework in DH and computational linguistics. Annotators can optionally add more theme and emotion tags. The interface also features a two-dimensional emotional coordinate based on the arousal-valence model (Russell 1980), a numeric affective representation previously tested in other cultural domains such as music (Yang et al. 2008). We expect future evaluation of its role in poetry collection building and discovery. To encourage thoughtful, grounded annotation, the interface requires annotators to provide brief explanations for their selected theme and emotion tags.

Figure 1. Poetry annotation interface (poem text blurred for copyright protection)

Workflow
Figure 2 illustrates the annotation workflows. Collections of poems are divided into small, adjustable Figure 2 illustrates the annotation workflows. Collections of poems are divided into small, adjustable worksets (e.g., twenty poems evenly sampled from the five cultural groups) to reduce bias from prior familiarity with specific groups. Worksets are automatically assigned upon completion, allowing asynchronous annotation with progress tracking. In the HTRC Data Capsule, which operates in offline mode, the workflow mirrors the online version, except that worksets are generated locally by an administrator and annotation results are stored locally. A separate administrative interface is available in the online version for user management, manual workset assignment, and export of annotation data.

Figure 2. Online and offline annotation workflows

Conclusion

We introduce Poetry Annotator, an open-source poetry annotation tool for annotating poems, including those in under-resourced languages, across online and offline environments. We hope other researchers will adapt this tool for their own annotation studies. In future work, we will use the annotated metadata to build a digital poetry library and analyze metadata and potential biases to promote equity.

References
  1. Abrami, G., Mehler, A., & Stoeckel, M. (2020, June). TextAnnotator: A web-based annotation suite for texts. DH2020.
  2. Bhyravajjula, S., Walsh, M., Preus, A., & Antoniak, M. (2025, November). so much depends/upon/a whitespace: Why Whitespace Matters for Poets and LLMs. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (pp. 35144–35161).
  3. Choi, K., & Kang, G. (2025). An analysis of poet demographic and thematic diversity in a poetry collection for inclusive AI. Information Research an international electronic journal, 30(iConf), 610-617.
  4. Colucci Cante, L., D’Angelo, S., Di Martino, B., & Graziano, M. (2024, July). Text annotation tools: A comprehensive review and comparative analysis. In International Conference on Complex, Intelligent, and Software Intensive Systems (pp. 353–362). Cham: Springer Nature Switzerland.
  5. Luka, M. E. & Leurs, K. (2020). Feminist Data Studies. In The International Encyclopedia of Gender, Media, and Communication (eds K. Ross, I. Bachmann, V. Cardo, S. Moorti, and M. Scarcelli). https://doi.org/10.1002/9781119429128.iegmc062
  6. Mauro, A. (2014). Small-Scale Big Data: Experimental Literature and Distributed Computing. Digital Humanities 2014, 263.
  7. Mohammad, S. M., & Turney, P. D. (2013). Crowdsourcing a word–emotion association lexicon. Computational intelligence, 29(3), 436–465.
  8. Nakayama, H., Kubo, T., Kamura, J., Taniguchi, Y., & Liang, X. (2018). doccano: Text Annotation Tool for Human. GitHub. https://github.com/doccano/doccano
  9. Russell, J. A. (1980). A circumplex model of affect. Journal of personality and social psychology, 39(6), 1161.
  10. Sprugnoli, R., Mambrini, F., Passarotti, M., & Moretti, G. (2023). The sentiment of Latin poetry. Annotation and automatic analysis of the Odes of Horace. IJCoL. Italian Journal of Computational Linguistics, 9(9-1).
  11. Stenetorp, P., Pyysalo, S., Topić, G., Ohta, T., Ananiadou, S., & Tsujii, J. (2012). Brat: a Web-based Tool for NLP-Assisted Text Annotation. In Proceedings of the Demonstrations Session at EACL 2012, 102–107. ACL Anthology.
  12. Yang, Y. H., Lin, Y. C., Su, Y. F., & Chen, H. H. (2008). A regression approach to music emotion recognition. IEEE Transactions on audio, speech, and language processing, 16(2), 448–457.
  13. Yogarajan, V., Dobbie, G., Pistotti, T., Bensemann, J., & Knowles, K. (2023). Challenges in annotating datasets to quantify bias in under-represented society. In CEUR Workshop Proceedings (Vol. 3547).