Daejeon, July 27–31
This study examines the system-level governance challenges arising from the large-scale use of financial news corpora in AI-driven sentiment analysis. As financial institutions increasingly rely on natural language processing (NLP) models for market prediction and risk assessment, news texts have evolved from mere information sources into high-frequency, high-value data assets embedded within critical information infrastructures. However, most existing sentiment analysis systems remain primarily optimized for model performance and computational efficiency, while copyright management and data governance are often treated as external or ex post legal concerns. As corpus usage scales across jurisdictions and diversified application contexts, this design separation gradually becomes a structural source of risk.
Recent international case law and regulatory developments indicate that the focus of copyright compliance assessment is shifting from users’ subjective intent to the system’s actual technical behavior and outputs. In the United States, Authors Guild v. Google (2015) recognized large-scale text digitization and full-text indexing as transformative fair use, with particular emphasis on whether the system functionally substitutes the original market (Authors Guild v. Google 2015). However, in Andy Warhol Foundation v. Goldsmith (2023), the U.S. Supreme Court significantly narrowed the scope of transformativeness, clarifying that even where form is altered, uses that closely align with the original work’s function and market position may still constitute infringement (Andy Warhol Foundation v. Goldsmith 2023). More recently, the ongoing litigation in The New York Times v. Microsoft & OpenAI (2023–) directly frames AI training processes and model outputs as technical behaviors capable of substituting the news market, highlighting that whether model outputs reproduce or replace protected expression has become a central legal concern (The New York Times v. Microsoft & OpenAI 2023–).
Outside the United States, Articles 3 and 4 of the EU Digital Single Market (DSM) Directive (2019) establish explicit exceptions for Text and Data Mining (TDM), while simultaneously allowing rightholders to reserve their rights through machine-readable opt-out mechanisms (European Parliament / Council 2019: Arts. 3–4). This regulatory design effectively requires corpus-processing systems to possess the technical capacity to detect, interpret, and enforce rights metadata. Although Japan’s Copyright Act adopts a comparatively broad information analysis exception under Article 30-4 (Japan 1970), recent AI policy guidelines issued by the Agency for Cultural Affairs increasingly emphasize purpose separation between training data and model outputs, in order to prevent the retrospective reconstruction of protected expression (Agency for Cultural Affairs 2024).
Taken together, these judicial and regulatory developments convey a clear signal: copyright compliance is no longer merely a matter of legal interpretation, but is highly dependent on concrete architectural choices in data access, transformation, storage, and output layers of corpus-processing systems.
From a systems perspective, this paper conceptualizes financial sentiment analysis as a multi-stage data processing pipeline encompassing corpus ingestion, preprocessing, and analytical deployment. Copyright-related risks do not arise at a single technical stage; rather, they accumulate as data undergo repeated copying, embedding, feature transformation, and output generation. In the absence of explicit governance mechanisms, system operators lack visibility into how specific corpora are used across training, validation, and production environments, resulting in opaque compliance boundaries and significantly increased coordination costs for cross-jurisdictional deployment (IBM 2024).
To address this structural problem, we propose a financial corpus licensing platform as a middleware layer embedded within NLP pipelines. Rather than functioning solely as a registry or licensing lookup service, the platform is designed as an institutional control plane that dynamically regulates data access rights, usage purposes, and accountability throughout the system lifecycle. In doing so, it translates the normative requirements articulated in international copyright regimes into executable system rules.
The proposed system architecture consists of five logical layers:
Within this architecture, licensing terms are no longer treated as external contractual obligations but are translated into control rules that take effect at runtime. For example, corpora licensed exclusively for model training may be delivered only in non-human-readable formats and blocked from deployment or external service endpoints, directly addressing the market substitution concerns articulated in Warhol and NYT v. OpenAI.
The contribution of this study lies in demonstrating how copyright governance can be reconceptualized as a system architecture and data governance problem. By introducing a licensing-aware middleware layer and clearly separating control and audit planes, this work proposes an actionable architectural pattern for embedding the constraints articulated by international case law and regulation into data-intensive AI systems. Beyond financial technology applications, the proposed design offers a transferable blueprint for sustainable corpus governance in other research and industrial domains that rely heavily on large-scale textual data.