Value-Action Alignment in Large Language Models under Privacy-Prosocial Conflict

71d ago · Global · primary source: export.arxiv.org

A new study from researcher Xiyang Hu introduces a protocol to test whether large language models act consistently with their stated privacy and prosocial values when those values conflict, revealing substantial variation across models [1]. The work, posted to arXiv on 7 January 2026 and revised on 28 June 2026, addresses a gap in how AI systems are evaluated when making decisions about personal data [1]. Existing assessments often measure privacy attitudes or sharing intentions separately, making it difficult to determine whether a model’s expressed values jointly predict its downstream actions as they do in human behavior [2]. Hu’s protocol administers standardized questionnaires for privacy attitudes, prosocialness, and acceptance of data sharing within a single, history-carrying session [2]. Multi-group structural equation modeling, or MGSEM, is then used to map the relationships from privacy concerns and prosocialness to data-sharing decisions [1]. The paper also proposes a new metric called the Value-Action Alignment Rate, or VAAR, which aggregates path-level evidence to produce a human-referenced directional agreement score [2]. Across the models tested, the researchers observed stable but model-specific profiles and what they describe as “substantial heterogeneity in value-action alignment” [1]. The study arrives as large language models are increasingly deployed to simulate human decision-making in sensitive domains, including personal data sharing, where privacy and prosocial motivations can push choices in opposite directions [2]. The findings suggest that a model’s stated attitudes do not uniformly translate into consistent data-sharing behaviors, a result that carries implications for developers and regulators weighing the use of AI in contexts that demand predictable, value-driven decisions. The submission, totaling 501 KB, lists Xiyang Hu as the corresponding author and is hosted on arXiv under the Computation and Language category [1]. The paper’s code and data are associated with the Hugging Face platform, according to the article’s supplementary links [1].

safety-researchresearch-paper

Background sources we checked (6)
  • arxiv.org ↗ Large language models (LLMs) are increasingly used to simulate decision-making tasks involving personal data sharing, where privacy concerns and prosocial motivations can push choices in opposite directions. Existing evaluations often measure privacy-related attitudes or sharing …
  • arxiv.org ↗ # A Universal Catalyst for First-Order Optimization ... arXiv (Cornell University), 2015. Preprint. 185 citations. ... We introduce a generic scheme for accelerating first-order optimization methods in the sense of Nesterov, which builds upon a new analysis of the accelerated pro…
  • arxiv.org ↗ CatalyzeX Code Finder for Papers (What is CatalyzeX?) ... DagsHub Toggle ... DagsHub (What is DagsHub?)…
  • arxiv.org ↗ CatalyzeX Code Finder for Papers (What is CatalyzeX?) ... DagsHub Toggle ... DagsHub (What is DagsHub?)…
  • en.wikipedia.org ↗ Sustainable Development Goals (abbr. SDGs) were adopted in 2015 by all United Nations (UN) members for the 2030 Agenda for Sustainable Development. The aim of the 17 global goals is "peace and prosperity for people and the planet", tackling climate change, and working to preserv…
  • en.wikipedia.org ↗ In molecular biology, a transcription factor (TF) (or sequence-specific DNA-binding factor) is a protein that controls the rate of transcription of genetic information from DNA to messenger RNA, by binding to DNA sequences. Specificity can be due to sequence motifs, or epigenetic…

Sources

Spot something wrong? Report an issue