DH 2026

Daejeon, July 27–31

Fri, July 3111:00–12:30S105108
Short Paper

Modular Architecture of a Financial Corpus Collective Licensing Platform

Chian-Yu Ye
National Chengchi University, Taiwan · 111601136@g.nccu.edu.tw
Chao-Lin Liu
National Chengchi University, Taiwan · chaolin@g.nccu.edu.tw
Zong-Ru Yang
National Chengchi University, Taiwan · 113753211@g.nccu.edu.tw

This study examines the system-level governance challenges arising from the large-scale use of financial news corpora in AI-driven sentiment analysis. As financial institutions increasingly rely on natural language processing (NLP) models for market prediction and risk assessment, news texts have evolved from mere information sources into high-frequency, high-value data assets embedded within critical information infrastructures. However, most existing sentiment analysis systems remain primarily optimized for model performance and computational efficiency, while copyright management and data governance are often treated as external or ex post legal concerns. As corpus usage scales across jurisdictions and diversified application contexts, this design separation gradually becomes a structural source of risk.

Recent international case law and regulatory developments indicate that the focus of copyright compliance assessment is shifting from users’ subjective intent to the system’s actual technical behavior and outputs. In the United States, Authors Guild v. Google (2015) recognized large-scale text digitization and full-text indexing as transformative fair use, with particular emphasis on whether the system functionally substitutes the original market (Authors Guild v. Google 2015). However, in Andy Warhol Foundation v. Goldsmith (2023), the U.S. Supreme Court significantly narrowed the scope of transformativeness, clarifying that even where form is altered, uses that closely align with the original work’s function and market position may still constitute infringement (Andy Warhol Foundation v. Goldsmith 2023). More recently, the ongoing litigation in The New York Times v. Microsoft & OpenAI (2023–) directly frames AI training processes and model outputs as technical behaviors capable of substituting the news market, highlighting that whether model outputs reproduce or replace protected expression has become a central legal concern (The New York Times v. Microsoft & OpenAI 2023–).

Outside the United States, Articles 3 and 4 of the EU Digital Single Market (DSM) Directive (2019) establish explicit exceptions for Text and Data Mining (TDM), while simultaneously allowing rightholders to reserve their rights through machine-readable opt-out mechanisms (European Parliament / Council 2019: Arts. 3–4). This regulatory design effectively requires corpus-processing systems to possess the technical capacity to detect, interpret, and enforce rights metadata. Although Japan’s Copyright Act adopts a comparatively broad information analysis exception under Article 30-4 (Japan 1970), recent AI policy guidelines issued by the Agency for Cultural Affairs increasingly emphasize purpose separation between training data and model outputs, in order to prevent the retrospective reconstruction of protected expression (Agency for Cultural Affairs 2024).

Taken together, these judicial and regulatory developments convey a clear signal: copyright compliance is no longer merely a matter of legal interpretation, but is highly dependent on concrete architectural choices in data access, transformation, storage, and output layers of corpus-processing systems.

From a systems perspective, this paper conceptualizes financial sentiment analysis as a multi-stage data processing pipeline encompassing corpus ingestion, preprocessing, and analytical deployment. Copyright-related risks do not arise at a single technical stage; rather, they accumulate as data undergo repeated copying, embedding, feature transformation, and output generation. In the absence of explicit governance mechanisms, system operators lack visibility into how specific corpora are used across training, validation, and production environments, resulting in opaque compliance boundaries and significantly increased coordination costs for cross-jurisdictional deployment (IBM 2024).

To address this structural problem, we propose a financial corpus licensing platform as a middleware layer embedded within NLP pipelines. Rather than functioning solely as a registry or licensing lookup service, the platform is designed as an institutional control plane that dynamically regulates data access rights, usage purposes, and accountability throughout the system lifecycle. In doing so, it translates the normative requirements articulated in international copyright regimes into executable system rules.

The proposed system architecture consists of five logical layers:

  • Ingestion Layer: News providers register corpus metadata through a rights registration service. The system generates persistent corpus identifiers linked to ownership information, applicable jurisdictions, licensing scope, and usage constraints (e.g., EU TDM opt-out markers).
  • Metadata & Policy Store: A centralized repository manages rights metadata, purpose tags, and jurisdiction-aware policy rules, serving as the basis for automated authorization decisions.
  • Access Control Layer (License-Aware API Gateway): All corpus access is mediated through license-aware APIs. Access requests must explicitly declare their intended purpose (e.g., model training, validation, or inference) and are evaluated against policy rules prior to data delivery.
  • Processing Layer (NLP Pipeline): Authorized corpora enter preprocessing and modeling workflows, with system-level enforcement of separation between raw texts and derived features (such as vector embeddings or sentiment indicators), thereby reducing the risk that model outputs reproduce protected expression.
  • Audit & Settlement Layer: All corpus usage events are recorded in machine-readable form, supporting compliance auditing, anomaly detection, and—where applicable—automated royalty distribution through smart contract mechanisms.

Within this architecture, licensing terms are no longer treated as external contractual obligations but are translated into control rules that take effect at runtime. For example, corpora licensed exclusively for model training may be delivered only in non-human-readable formats and blocked from deployment or external service endpoints, directly addressing the market substitution concerns articulated in Warhol and NYT v. OpenAI.

The contribution of this study lies in demonstrating how copyright governance can be reconceptualized as a system architecture and data governance problem. By introducing a licensing-aware middleware layer and clearly separating control and audit planes, this work proposes an actionable architectural pattern for embedding the constraints articulated by international case law and regulation into data-intensive AI systems. Beyond financial technology applications, the proposed design offers a transferable blueprint for sustainable corpus governance in other research and industrial domains that rely heavily on large-scale textual data.

References
  1. Agency for Cultural Affairs (Japan) (2024): General Understanding on AI and Copyright in Japan. Tokyo: Agency for Cultural Affairs, Government of Japan.
  2. Andy Warhol Foundation for the Visual Arts, Inc. v. Goldsmith (2023): 598 U.S. 508. United States Supreme Court.
  3. Authors Guild, Inc. v. Google, Inc. (2015): 804 F.3d 202. United States Court of Appeals for the Second Circuit.
  4. European Parliament / Council of the European Union (2019): "Directive (EU) 2019/790 of the European Parliament and of the Council of 17 April 2019 on copyright and related rights in the Digital Single Market and amending Directives 96/9/EC and 2001/29/EC", in: Official Journal of the European Union L 130: 92–125.
  5. Japan (1970): Copyright Act (Act No. 48 of 6 May 1970, as amended), Article 30-4. Tokyo: Government of Japan.
  6. The New York Times Company v. Microsoft Corporation and OpenAI, Inc. (2023–): Case No. 1:23-cv-11195. United States District Court for the Southern District of New York, filed 27 December 2023.
  7. IBM (2024): "Impact on Data Governance with Generative AI – Part One" https://www.ibm.com/think/insights/impact-on-data-governance-with-generative-ai-part-one.
  8. IBM (2024): "Impact on Data Governance with Generative AI – Part One" https://www.ibm.com/think/insights/impact-on-data-governance-with-generative-ai-part-one