Score: 0

VisKnow: Constructing Visual Knowledge Base for Object Understanding

Published: December 9, 2025 | arXiv ID: 2512.08221v1

By: Ziwei Yao , Qiyang Wan , Ruiping Wang and more

Potential Business Impact:

Helps computers understand what things are, not just names.

Business Areas:

Image Recognition Data and Analytics, Software

Understanding objects is fundamental to computer vision. Beyond object recognition that provides only a category label as typical output, in-depth object understanding represents a comprehensive perception of an object category, involving its components, appearance characteristics, inter-category relationships, contextual background knowledge, etc. Developing such capability requires sufficient multi-modal data, including visual annotations such as parts, attributes, and co-occurrences for specific tasks, as well as textual knowledge to support high-level tasks like reasoning and question answering. However, these data are generally task-oriented and not systematically organized enough to achieve the expected understanding of object categories. In response, we propose the Visual Knowledge Base that structures multi-modal object knowledge as graphs, and present a construction framework named VisKnow that extracts multi-modal, object-level knowledge for object understanding. This framework integrates enriched aligned text and image-source knowledge with region annotations at both object and part levels through a combination of expert design and large-scale model application. As a specific case study, we construct AnimalKB, a structured animal knowledge base covering 406 animal categories, which contains 22K textual knowledge triplets extracted from encyclopedic documents, 420K images, and corresponding region annotations. A series of experiments showcase how AnimalKB enhances object-level visual tasks such as zero-shot recognition and fine-grained VQA, and serves as challenging benchmarks for knowledge graph completion and part segmentation. Our findings highlight the potential of automatically constructing visual knowledge bases to advance visual understanding and its practical applications. The project page is available at https://vipl-vsu.github.io/VisKnow.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

CV and Pattern Recognition

Teaches computers to understand how the world works.

25 Nov 2025 2

88%

Seeing and Knowing in the Wild: Open-domain Visual Entity Recognition with Large-scale Knowledge Graphs via Contrastive Learning

CV and Pattern Recognition

Helps computers recognize any object in pictures.

15 Oct 2025 1

88%

A Comprehensive Survey of Knowledge-Based Vision Question Answering Systems: The Lifecycle of Knowledge in Visual Reasoning Task

CV and Pattern Recognition

Helps computers answer questions using pictures and facts.

24 Apr 2025 0

View PDF Login to Bookmark

Page Count

36 pages

VisKnow: Constructing Visual Knowledge Base for Object Understanding

Helps computers understand what things are, not just names.

Technical Abstract

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

Seeing and Knowing in the Wild: Open-domain Visual Entity Recognition with Large-scale Knowledge Graphs via Contrastive Learning

A Comprehensive Survey of Knowledge-Based Vision Question Answering Systems: The Lifecycle of Knowledge in Visual Reasoning Task