Creating Multilingual Mental Health Dialogue Datasets: Limits of Persona-Based Localization via Nationality and Language
- company Hugging Face
- company arXivLabs
- location Bengali
- location California
- location Hindi
- location Mandarin
- location Taiwan
- model LLMs
Researchers have found that simply swapping nationality and language parameters in synthetic personas is insufficient for creating reliable multilingual mental health datasets, exposing clinical inconsistencies when large language models assess depression severity across Mandarin, Bengali, and Hindi. The study, submitted on 17 Jun 2026, investigated whether persona-based localization methods, commonly used to generate English-centric clinical training data, could be extended to non-English contexts [1]. The team generated clinical dialogues in Mandarin, Bengali, and Hindi by modifying nationality and language parameters in existing personas, then tested how different large language models (LLMs) evaluated the depression severity of these datasets against an English baseline [2]. The findings showed that the approach can introduce clinical inconsistency across languages, and that LLM judge models often exhibit inaccuracies when assessing depression severity in non-English texts, with performance varying across different models [2]. The work highlights a systemic limitation in current data-generation practices. Most validated personas rely on English-centric contexts, and the paper argues that merely translating demographic tags does not capture the cultural framing of mental health symptoms [2]. The authors call for culturally responsive data generation to ensure equitable mental health systems globally [2]. LLMs are a type of machine learning model designed for natural language processing tasks such as language generation, trained with self-supervised learning on vast amounts of text [7]. Their application in mental health has drawn interest because of a critical shortage of high-quality datasets for training and evaluating digital support systems [2]. The arXiv paper appears in the Computation and Language category, and its abstract page includes links to code and data repositories through the arXivLabs integration with Hugging Face Spaces, a collaboration that allows researchers to embed interactive demos directly alongside papers [4][5]. The broader AI landscape has seen rapid shifts in model development costs and accessibility. Chinese firm DeepSeek, founded in July 2023, reported training its V3 model for US$6 million, a fraction of the estimated US$100 million cost for OpenAI's GPT-4 in 2023 [6]. DeepSeek's models are described as open-weight, with parameters openly shared under free and open-source software licenses, though training data is not openly licensed [6]. The company recruits from top Chinese universities and hires outside traditional computer science fields to broaden its models' knowledge [6]. Hugging Face's arXiv integration, launched in November 2022, was designed to make papers more accessible by linking to open-source demos built with tools such as Gradio and Streamlit [3]. The feature allows users to try models without writing code, and the organization has stated that demos help a wider audience identify and debug biases and other issues [3].
research-paper
Background sources we checked (7)
- arxiv.org ↗ AI and large language models (LLMs) have emerged as promising tools to address global mental health challenges. Despite the global nature of these challenges, there remains a critical shortage of high-quality datasets for training and evaluating such systems. To mitigate this gap…
- huggingface.co ↗ Hugging Face Machine Learning Demos on arXiv Back to Articles ... # Hugging Face Machine Learning Demos on arXiv Published November 17, 2022 Update on GitHub Upvote 1 - - - - - Abubakar Abid abidlabs Follow …
- info.arxiv.org ↗ ## Hugging Face Spaces ... Hugging Face code repositories, About Hugging Face ... Collaborators: Abubakar Abid, Omar Sanseviero, Ahsen Khaliq, and the Hugging Face team ... Hugging Face Spaces includes links to demos created by the community or the authors themselves. By going to…
- huggingface.co ↗ Demos on Hugging Face Spaces allow a wide audience to try out state-of-the-art machine learning research without writing any code. Hugging Face and ArXiv have collaborated to embed these demos directly along side papers on ArXiv! ... Thanks to this integration, users can now find…
- en.wikipedia.org ↗ Hangzhou DeepSeek Artificial Intelligence Basic Technology Research Co., Ltd., doing business as DeepSeek, is a Chinese artificial intelligence (AI) company that develops large language models (LLMs). Based in Hangzhou, Zhejiang, DeepSeek is owned and funded by High-Flyer, a Chin…
- en.wikipedia.org ↗ A large language model (LLM) is a type of machine learning model designed for natural language processing tasks such as language generation. LLMs are language models with many parameters, and are trained with self-supervised learning on a vast amount of text.…
- en.wikipedia.org ↗ Douwe Kiela is a Dutch-American research scientist and entrepreneur working in the field of artificial intelligence with a focus on machine learning and natural language processing. He is a research scientist director at Google DeepMind. He previously co-founded and served as CEO…