Score: 1

A Scalable Unsupervised Framework for multi-aspect labeling of Multilingual and Multi-Domain Review Data

Published: May 14, 2025 | arXiv ID: 2505.09286v1

By: Jiin Park, Misuk Kim

Potential Business Impact:

Helps computers understand reviews in any language.

Business Areas:

Natural Language Processing Artificial Intelligence, Data and Analytics, Software

Effectively analyzing online review data is essential across industries. However, many existing studies are limited to specific domains and languages or depend on supervised learning approaches that require large-scale labeled datasets. To address these limitations, we propose a multilingual, scalable, and unsupervised framework for cross-domain aspect detection. This framework is designed for multi-aspect labeling of multilingual and multi-domain review data. In this study, we apply automatic labeling to Korean and English review datasets spanning various domains and assess the quality of the generated labels through extensive experiments. Aspect category candidates are first extracted through clustering, and each review is then represented as an aspect-aware embedding vector using negative sampling. To evaluate the framework, we conduct multi-aspect labeling and fine-tune several pretrained language models to measure the effectiveness of the automatically generated labels. Results show that these models achieve high performance, demonstrating that the labels are suitable for training. Furthermore, comparisons with publicly available large language models highlight the framework's superior consistency and scalability when processing large-scale data. A human evaluation also confirms that the quality of the automatic labels is comparable to those created manually. This study demonstrates the potential of a robust multi-aspect labeling approach that overcomes limitations of supervised methods and is adaptable to multilingual, multi-domain environments. Future research will explore automatic review summarization and the integration of artificial intelligence agents to further improve the efficiency and depth of review analysis.

Multi-domain Multilingual Sentiment Analysis in Industry: Predicting Aspect-based Opinion Quadruples

Computation and Language

Helps computers understand opinions about things.

15 May 2025 1

88%

Context-Aware Hierarchical Taxonomy Generation for Scientific Papers via LLM-Guided Multi-Aspect Clustering

Computation and Language

Organizes science papers into helpful, detailed lists.

23 Sep 2025 1

88%

From Annotation to Adaptation: Metrics, Synthetic Data, and Aspect Extraction for Aspect-Based Sentiment Analysis with Large Language Models

Computation and Language

Helps computers understand feelings about sports.

26 Mar 2025 1

View PDF Login to Bookmark

Repos / Data Links

github.com github.com

Page Count

36 pages

A Scalable Unsupervised Framework for multi-aspect labeling of Multilingual and Multi-Domain Review Data

Helps computers understand reviews in any language.

Technical Abstract

Multi-domain Multilingual Sentiment Analysis in Industry: Predicting Aspect-based Opinion Quadruples

Context-Aware Hierarchical Taxonomy Generation for Scientific Papers via LLM-Guided Multi-Aspect Clustering

From Annotation to Adaptation: Metrics, Synthetic Data, and Aspect Extraction for Aspect-Based Sentiment Analysis with Large Language Models