DH 2026

Daejeon, July 27–31

Wed, July 2909:00–10:30S052209-211
Short Paper

Remembering the Redditor: A Careful Approach to 'Big-But-Intimate Data' on Advice Forums

Joseph Reagle
Northeastern University, United States of America · j.reagle@neu.edu


Reddit is the most vital manifestation of the centuries-old advice genre. Subreddits such as r/relationship_advice and r/AmItheAsshole are repositories of the intimate conundrums of tens of thousands of people, discussed by hundreds of thousands, and seen by an audience of millions. Using such big-but-intimate data has methodological and ethical challenges ( Reddit is the most vital manifestation of the centuries-old advice genre. Subreddits such as r/relationship_advice and r/AmItheAsshole are repositories of the intimate conundrums of tens of thousands of people, discussed by hundreds of thousands, and seen by an audience of millions. Using such big-but-intimate data has methodological and ethical challenges ( Graham / Milligan / Weingart 2015; Milligan 2019) that scholars of traditional print advice columns never encountered. While most advice genre scholarship and Reddit research presume “the data is already public,” this stance often fails to “remember the human” behind online disclosures ( Zimmer 2010; Fiesler et al. 2024). Additionally, digital humanists span disciplines (e.g., history and communication studies) and produce reports across genres (e.g., journal articles, journalistic essays, and scholarly and trade books). These disciplines and types of reports vary in how they use and excerpt online sources. As part of my research on the advice genre and Reddit, I developed an approach for care in selection and reporting (CSR) of online messages.

People often share disclosures in public online forums that would be sensitive in other contexts (e.g., information about their health, work, and relationships). I describe such content as public quasi-sensitive disclosures, hereafter disclosures. Though these disclosures are not sensitive in a regulatory or Ethical/Institutional Review Board sense, the reporting of such disclosures merits care. Even though researchers might have no interactions with online sources, and consequently there are no research subjects, there are some risks. Consequential harm follows when third parties use disclosures reported by researchers to sources’ detriment. For example, a report that includes disclosures of taboo or illegal activities could prompt action from government agencies, employers, family, or friends. Though such harm is possible, I’m not aware of it befalling an online source that appeared in a report. The more probable affective harm is the negative feelings a source has when they see their messages reported and discussed outside the original context ( Nissenbaum 2004: 119; Klassen / Fiesler 2022). For example, a user discussing a failed relationship might be annoyed to see the post excerpted and used in an article or book.

Taking care in selecting and reporting public quasi-sensitive disclosures is a middle path. It plots a course between the position that all such data is public and can be used without constraint and the position that consent must be obtained from all sources. The CSR approach attempts to lessen the risks of reporting online disclosures. It complements ethically disguising sources ( Bruckman 2002), an important technique that is often error-prone in practice ( Reagle 2022).

The CSR framework attends to how a researcher selects, presents, and tracks sources in their manuscripts. This how is contingent on why a researcher is using online disclosures. For example, I was writing about the advice genre and subreddits. I wanted to describe the history of the forums (e.g., a subreddit’s creation), how sources conceived of their participation (e.g., wanting to be of help), and associated issues (e.g., pseudonymous expertise). I also wanted to provide a few illustrative conundrums. The framework asks the researcher to consider these questions:

  • how dated is the message?
  • is the identity of the original poster (OP) a pseudonym, throwaway, or deleted account?
  • is the gravity of the disclosure mild, moderate, or major given the level and duration of expressed and potential distress? ( Wallace et al. 2023).
  • did the OP engage with the audience via comments and updates?
  • is the message already “internet-famous” and distributed off-site via syndication?

For example, a disclosure of major gravity might be avoided or only alluded to without excerpt or citation. A disclosure of moderate gravity and not widely distributed across the web might be disguised in both source and prose. An eleven-year-old disclosure, posted by a throwaway and widely distributed and commented on across the web, might be excerpted and cited. A disclosure whose prose would be significantly excerpted might warrant asking the OP for objection or consent.

I manage disclosures within a manuscript with annotations that track the provenance and treatment of each source. For example, in my markdown manuscripts, I embed structured HTML comments of the form: <!-- ethics: c:[[case]] d:[[date]] g:[[gravity]] i:[[identity]] p:[[prose]] s:[[syndicated]] u:[[updated]] r:[[responded]] n:[[note]] --> A relationship-advice post about a mild conflict during a camping trip, posted eleven years ago, by a throwaway account, excerpts of which I use verbatim, which was syndicated off-site, was updated by the OP, and whom I did not contact would look like: <!-- ethics: c:camping d:old g:mild i:throwaway p:verbatim s:yes u:yes r:na n:na -->. These comments do not appear in the finished manuscript, but they are easy for me to track and analyze. In my bibliographic database, those sources I disguise have both a proper entry and a corresponding disguised citation yielding undated citations in the manuscript, such as “Disguised_OP, ‘Gaslighting (Disguised Post),’ Reddit.”

This careful selection and reporting of sources offers researchers in the digital humanities a way of using public quasi-sensitive disclosures. It offers a middle path between the extremes of wantonly reporting disclosures and being inappropriately encumbered by the ethical requirements rightly reserved for actual human-subjects research.

References
  1. Bruckman, Amy (2002): " Studying the amateur artist: a perspective on disguising data collected in human subjects research on the Internet", in: Ethics and Information Technology 4 (3):
  2. Fiesler, Casey / Zimmer, Michael / Proferes, Nicholas / Gilbert, Sarah / Jones, Naiyan (2024): "Remember the human: A systematic review of ethical considerations in Reddit research", in: Proceedings of the ACM on Human-Computer Interaction 8 (GROUP): 10.1145/3633070.
  3. Graham, Shawn / Milligan, Ian / Weingart, Scott (2015): Exploring big historical data: The historian’s macroscope. Hackensack, NJ: Imperial College Press.
  4. Klassen, Shamika / Fiesler, Casey (2022): " “This Isn’t your data, Friend”: Black twitter as a case study on research ethics for public data", in: Social Media + Society 8 (4):
  5. Milligan, Ian (2019): History in the age of abundance?: How the web is transforming historical research. Montreal: McGill-Queen’s University Press.
  6. Nissenbaum, Helen (2004): " Privacy as contextual integrity", in: Washington Wall Review 79: 101–139.
  7. Reagle, Joseph (2022): "Disguising Reddit sources and the efficacy of ethical research", in: Ethics and Information Technology 24 (3): 10.1007/s10676-022-09663-w.
  8. Wallace, Denise / Cooper, Nicholas R. / Sel, Alejandra / Russo, Riccardo (2023): " The social readjustment rating scale: Updated and modernised", in: PLOS ONE 18 (12): e0295943.
  9. Zimmer, Michael (2010): "„But the data is already public“: On the ethics of research in Facebook", in: Ethics and Information Technology 12 (4): 10.1007/s10676-010-9227-5.