Related papers: Diff-Oracle: Deciphering Oracle Bone Scripts with Controllable Diffusion Model

Diff-Oracle: Deciphering Oracle Bone Scripts with Controllable Diffusion Model

URL: http://arxiv.org/abs/2312.13631v2
Date: Mon, 8 Jul 2024 07:34:29 GMT
Title: Diff-Oracle: Deciphering Oracle Bone Scripts with Controllable Diffusion Model
Authors: Jing Li, Qiu-Feng Wang, Siyuan Wang, Rui Zhang, Kaizhu Huang, Erik Cambria,
Abstract summary: Deciphering oracle bone scripts plays an important role in Chinese archaeology and philology. Diff-Oracle is a novel approach based on diffusion models to generate controllable oracle characters. Diff-Oracle substantially benefits downstream oracle character recognition, outperforming all existing SOTAs by a large margin.
Score: 48.956844881630886
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Abstract: Deciphering oracle bone scripts plays an important role in Chinese archaeology and philology. However, a significant challenge remains due to the scarcity of oracle character images. To overcome this issue, we propose Diff-Oracle, a novel approach based on diffusion models to generate a diverse range of controllable oracle characters. Unlike traditional diffusion models that operate primarily on text prompts, Diff-Oracle incorporates a style encoder that utilizes style reference images to control the generation style. This encoder extracts style prompts from existing oracle character images, where style details are converted into a text embedding format via a pretrained language-vision model. On the other hand, a content encoder is integrated within Diff-Oracle to capture specific content details from content reference images, ensuring that the generated characters accurately represent the intended glyphs. To effectively train Diff-Oracle, we pre-generate pixel-level paired oracle character images (i.e., style and content images) by an image-to-image translation model. Extensive qualitative and quantitative experiments are conducted on datasets Oracle-241 and OBC306. While significantly surpassing present generative methods in terms of image generation, Diff-Oracle substantially benefits downstream oracle character recognition, outperforming all existing SOTAs by a large margin. In particular, on the challenging OBC306 dataset, Diff-Oracle leads to an accuracy gain of 7.70% in the zero-shot setting and is able to recognize unseen oracle character images with the accuracy of 84.62%, achieving a new benchmark for deciphering oracle bone scripts.

Related papers

OracleFusion: Assisting the Decipherment of Oracle Bone Script with Structurally Constrained Semantic Typography [58.790901822971094]
Oracle Bone Script (OBS) encapsulates the cultural records and intellectual expressions of ancient civilizations.<n>Despite the discovery of approximately 4,500 OBS characters, only about 1,600 have been deciphered.<n>This paper proposes a novel two-stage semantic framework, named OracleFusion.
arXiv Detail & Related papers (2025-06-26T08:56:07Z)
BD-Diff: Generative Diffusion Model for Image Deblurring on Unknown Domains with Blur-Decoupled Learning [55.21345354747609]
BD-Diff is a generative-diffusion-based model designed to enhance deblurring performance on unknown domains. We employ two Q-Formers as structural representations and blur patterns extractors separately. We introduce a reconstruction task to make the structural features and blur patterns complementary.
arXiv Detail & Related papers (2025-02-03T17:00:40Z)
OracleSage: Towards Unified Visual-Linguistic Understanding of Oracle Bone Scripts through Cross-Modal Knowledge Fusion [19.788896054132053]
Oracle bone script (OBS), as China's earliest mature writing system, present significant challenges in automatic recognition. We introduce OracleSage, a novel cross-modal framework that integrates hierarchical visual understanding with graph-based semantic reasoning.
arXiv Detail & Related papers (2024-11-26T19:26:06Z)
Storynizor: Consistent Story Generation via Inter-Frame Synchronized and Shuffled ID Injection [27.412361280397057]
We introduce Storynizor, a model capable of generating coherent stories with strong inter-frame character consistency. Key innovation of Storynizor lies in its key modules: ID-Synchronizer and ID-Injector. To facilitate the training of Storynizor, we have curated a novel dataset called StoryDB comprising 100, 000 images.
arXiv Detail & Related papers (2024-09-29T09:15:51Z)
Unsupervised Attention Regularization Based Domain Adaptation for Oracle Character Recognition [59.05212866862219]
The study of oracle characters plays an important role in Chinese archaeology and philology. The difficulty of collecting and annotating real-world scanned oracle characters hinders the development of oracle character recognition. We develop a novel unsupervised domain adaptation (UDA) method to transfer recognition knowledge from labeled handprinted oracle characters to unlabeled scanned data.
arXiv Detail & Related papers (2024-09-24T09:07:05Z)
Openstory++: A Large-scale Dataset and Benchmark for Instance-aware Open-domain Visual Storytelling [81.69474860607542]
We present Openstory++, a large-scale dataset combining additional instance-level annotations with both images and text. We also present Cohere-Bench, a pioneering benchmark framework for evaluating the image generation tasks when long multimodal context is provided.
arXiv Detail & Related papers (2024-08-07T11:20:37Z)
Heracles: A Hybrid SSM-Transformer Model for High-Resolution Image and Time-Series Analysis [23.511807886483087]
Heracles is a novel SSM that integrates a local SSM, a global SSM, and an attention-based token interaction module. Heracles achieves state-of-the-art performance on the ImageNet dataset with 84.5% top-1 accuracy. Heracles excels in transfer learning tasks on datasets such as CIFAR-10, CIFAR-100, Oxford Flowers, and Stanford Cars.
arXiv Detail & Related papers (2024-03-26T19:29:21Z)
Scaling Autoregressive Models for Content-Rich Text-to-Image Generation [95.02406834386814]
Parti treats text-to-image generation as a sequence-to-sequence modeling problem. Parti uses a Transformer-based image tokenizer, ViT-VQGAN, to encode images as sequences of discrete tokens. PartiPrompts (P2) is a new holistic benchmark of over 1600 English prompts.
arXiv Detail & Related papers (2022-06-22T01:11:29Z)
Oracle-MNIST: a Realistic Image Dataset for Benchmarking Machine Learning Algorithms [57.29464116557734]
We introduce the Oracle-MNIST dataset, comprising of 28$times $28 grayscale images of 30,222 ancient characters. The training set totally consists of 27,222 images, and the test set contains 300 images per class.
arXiv Detail & Related papers (2022-05-19T09:57:45Z)
Unsupervised Structure-Texture Separation Network for Oracle Character Recognition [70.29024469395608]
Oracle bone script is the earliest-known Chinese writing system of the Shang dynasty and is precious to archeology and philology. We propose a structure-texture separation network (STSN), which is an end-to-end learning framework for joint disentanglement, transformation, adaptation and recognition.
arXiv Detail & Related papers (2022-05-13T10:27:02Z)

This list is automatically generated from the titles and abstracts of the papers in this site.

This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.