Score: 0

Evaluating LLM Alignment on Personality Inference from Real-World Interview Data

Published: September 16, 2025 | arXiv ID: 2509.13244v1

By: Jianfeng Zhu , Julina Maharjan , Xinyu Li and more

Potential Business Impact:

Computers can't guess your personality from talking.

Business Areas:

Natural Language Processing Artificial Intelligence, Data and Analytics, Software

Large Language Models (LLMs) are increasingly deployed in roles requiring nuanced psychological understanding, such as emotional support agents, counselors, and decision-making assistants. However, their ability to interpret human personality traits, a critical aspect of such applications, remains unexplored, particularly in ecologically valid conversational settings. While prior work has simulated LLM "personas" using discrete Big Five labels on social media data, the alignment of LLMs with continuous, ground-truth personality assessments derived from natural interactions is largely unexamined. To address this gap, we introduce a novel benchmark comprising semi-structured interview transcripts paired with validated continuous Big Five trait scores. Using this dataset, we systematically evaluate LLM performance across three paradigms: (1) zero-shot and chain-of-thought prompting with GPT-4.1 Mini, (2) LoRA-based fine-tuning applied to both RoBERTa and Meta-LLaMA architectures, and (3) regression using static embeddings from pretrained BERT and OpenAI's text-embedding-3-small. Our results reveal that all Pearson correlations between model predictions and ground-truth personality traits remain below 0.26, highlighting the limited alignment of current LLMs with validated psychological constructs. Chain-of-thought prompting offers minimal gains over zero-shot, suggesting that personality inference relies more on latent semantic representation than explicit reasoning. These findings underscore the challenges of aligning LLMs with complex human attributes and motivate future work on trait-specific prompting, context-aware modeling, and alignment-oriented fine-tuning.

From Five Dimensions to Many: Large Language Models as Precise and Interpretable Psychological Profilers

Artificial Intelligence

Computers guess your personality from a few answers.

5 Nov 2025 0

92%

Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans

Computation and Language

AI learns to argue and negotiate like people.

19 Sep 2025 1

92%

Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented Tasks

Human-Computer Interaction

Makes chatbots more likable with just enough personality.

11 Sep 2025 0

View PDF Login to Bookmark

Page Count

9 pages

Evaluating LLM Alignment on Personality Inference from Real-World Interview Data

Computers can't guess your personality from talking.

Technical Abstract

From Five Dimensions to Many: Large Language Models as Precise and Interpretable Psychological Profilers

Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans

Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented Tasks