Daejeon, July 27–31
A Pairwise Comparison Study of Critical Acclaim and Commercial Success in LLM Film Preference
Cultural taste has long been understood as a mechanism of social stratification. Bourdieu's Distinction (1984) set forth the foundational account of how aesthetic preferences reproduce and legitimize social hierarchy. Large language models (LLMs) introduce a new layer to this logic. Recommendation systems on streaming and social media platforms already shape cultural visibility through algorithmic curation (Alexander 2016; Seaver 2022). Yet, LLMs arguably extend this further into contexts where evaluative defaults are less visible (e.g., writing assistance, educational tools) and thus less anticipated. Recent work contends that LLMs reflect dominant patterns in training data and amplify convergence as users rely on the same systems across contexts (Sourati et al. 2026). Whereas that argument concerns output homogenization, this paper addresses the prior question of whether LLMs display preferences capable of driving such convergence in cultural domains. Before asking what LLM use does to human taste, we need to determine whether LLMs have something resembling "taste" of their own.
Instead of attempting to operationalize the Bourdieusian distinction between elite and popular taste, we isolate a measurable aspect of symbolic legitimacy: the tension between critical acclaim and commercial success. Operationalizing class-inflected distinctions would require ground-truth data on what cultural objects people across different socioeconomic positions actually consume — data that does not exist at the scale or granularity this type of study demands. We posit that the critical/commercial axis is one of the empirically tractable dimensions of cultural stratification available, and that systematic LLM preferences along this axis, if they exist, would be meaningful evidence that models encode something structurally analogous to the logic Bourdieu described.
Our corpus comprises 200 films partitioned into three sets based on their position within two ranking systems — the They Shoot Pictures, Don't They (TSPDT) top-1000, a meta-canonical index, and the Box Office Mojo all-time top-1000 by worldwide gross revenue. Set A (n=40) consists of films appearing in both lists, indicative of both critical and commercial recognition. Set B (n = 80) contains films on the TSPDT list that are absent from commercial rankings. Conversely, Set C (n = 80) contains films on commercial rankings that are absent from the TSPDT list. The corpus is sampled to include diverse languages and eras so that cultural context has less of an impact on set membership.
Preferences are elicited through a Bradley-Terry paired comparison design, a probabilistic model that infers latent preference orderings from binary choices (Bradley / Terry 1952; DiGiuseppe / Flynn 2026). Each model receives pairwise comparisons under a single prompt ("Which of these two films do you prefer?") via API, with responses constrained at temperature zero to isolate stable signals from stochastic variation. We conducted 4,000 comparisons per run across five independent runs per model — 20,000 comparisons per model in total — and fitted a Bradley-Terry maximum likelihood model to each run. We then aggregated the iterations to produce stable per-film log-strength estimates (λ). Eight models across four families were evaluated: Claude Haiku 4.5 and Sonnet 4.6 (Anthropic), GPT-5.4 Nano and GPT-5.4 (OpenAI), Qwen 2.5 Turbo and Plus (Alibaba), and Mistral Small 3.2 and Large 3 (Mistral). The two-tier, within-family design partially holds training lineage constant while varying model scale in order to enable scale comparisons.
To confirm that the captured preferences were not artifacts of stochastic noise, we verified the design against three reliability criteria: internal consistency at the item level (H1a; > 0.95), ranking stability across independent iterations (H1b; Spearman's 𝜌 = 0.831-0.925), and prompt-frame invariance (H1c), where rankings remained stable across semantically equivalent yet lexically distinct formulations (e.g., "Which do you prefer?" vs. "Which do you like?" vs. "Which is closer to your taste?"). These results suggest that the observed outputs reflect persistent ranking tendencies rather than fleeting surface-level cues.
The empirical results support H2, as all eight models exhibit what we term "critical acclaim orientation," operationalized as films appearing exclusively in the critical canon (Set B) winning substantially more than 50% of head-to-head matchups against commercially successful films (Set C). Win rates ranged from 65.6% to 87.8% across models (all p < .001), indicative of a robust preference for canonically prestigious films over commercially dominant ones. This pattern appears across all model families and capability tiers.
The relationship between dual-recognition films (Set A) and single-criterion films further clarifies the structure of this evaluative orientation. While Set A films strongly dominate the "popular-only" list (Set C) across all models with win rates ranging from 76.6% to 91.4%, they exhibit no consistent dominance over the "critical-only" list (Set B). In five of the eight evaluated models, comparisons between Set A and Set B failed to reach statistical significance. Consequently, H3 (predicting that dual-legitimacy films would defeat both critical-only and popular-only sets) is rejected. This suggests that critical prestige functions as the dominant evaluative signal, whereas commercial success contributes comparatively little independent influence once critical legitimacy is present.
The results also indicate that critical acclaim orientation intensifies with model scale (H4). Within each model family, larger-capability models displayed stronger preference differentials favoring critically canonized films over commercially successful non-canonical films. While the win rate for dual-legitimacy films (Set A) against the popular-only list (Set C) remains stable across tiers (small: 76.6%-89.7% vs. large: 85.2%-91.4%), the win rate for critical-only films (Set B) against Set C rises with model scale. One plausible interpretation is that larger models possess broader representational coverage of canonically discussed cinema, including films with lower popular visibility but substantial critical discourse density. Under this interpretation, scaling expands the model's ability to recognize and reproduce prestige signals embedded in textual training distributions.
These findings carry implications that cut across disciplinary lines. For digital humanities and cultural sociology, this study shifts the focus of algorithmic culture from representation to selection. Prior work has shown that evaluative hierarchies are latent in the distributional structure of word embeddings (Kozlowski et al. 2019); this study asks whether those hierarchies manifest in model outputs. The answer is yes, and this opens a broader methodological proposition. LLMs trained on the accumulated output of online cultural discourse may function as novel instruments for studying symbolic hierarchy. For AI ethics researchers, the findings expose a dimension of model behavior that the dominant bias framework misses. Most contemporary auditing approaches focus on demographic representation and discriminatory harms across social groups. Evaluative hierarchies in aesthetic domains pose a distinct challenge as there is no neutral baseline distribution against which "bias" can be measured. The choice is not between a biased model and a less biased one; it is between different value systems, none of which can claim neutrality. This study is intended as an opening provocation for that conversation, and a proof of concept that the empirical tools for having it exist.