AnyText: Multilingual Visual Text Generation And Editing

📌News

[2023.12.28] - Online demo is available here!
[2023.12.27] - 🧨We released the latest checkpoint(v1.1) and inference code, check on modelscope in Chinese.
[2023.12.05] - The paper is available at here.

💡Methodology

AnyText comprises a diffusion pipeline with two primary elements: an auxiliary latent module and a text embedding module. The former uses inputs like text glyph, position, and masked image to generate latent features for text generation or editing. The latter employs an OCR model for encoding stroke data as embeddings, which blend with image caption embeddings from the tokenizer to generate texts that seamlessly integrate with the background. We employed text-control diffusion loss and text perceptual loss for training to further enhance writing accuracy.

🛠Installation

# Install git (skip if already done)
conda install -c anaconda git
# Clone anytext code
git clone https://github.com/tyxsspa/AnyText.git
cd AnyText
# Prepare a font file; Arial Unicode MS is recommended, **you need to download it on your own**
mv your/path/to/arialuni.ttf ./font/Arial_Unicode.ttf
# Create a new environment and install packages as follows:
conda env create -f environment.yaml
conda activate anytext

🔮Inference

[Recommend]： We release a demo on ModelScope!

AnyText include two modes: Text Generation and Text Editing. Running the simple code below to perform inference in both modes and verify whether the environment is correctly installed.

python inference.py

If you have advanced GPU (with at least 20G memory), it is recommended to deploy our demo as below, which includes usage instruction, user interface and abundant examples.

python demo.py

Please note that when executing inference for the first time, the model files will be downloaded to: ~/.cache/modelscope/hub. If you need to modify the download directory, you can manually specify the environment variable: MODELSCOPE_CACHE.

🌄Gallery

📈Evaluation

We use Sentence Accuracy (Sen. ACC) and Normalized Edit Distance (NED) to evaluate the accuracy of generated text, and use the FID metric to assess the quality of generated images. Compared to existing methods, AnyText has a significant advantage in both Chinese and English text generation.

⏰TODOs

Release the model and inference code
Provide publicly accessible demo link
Release tools for merging weights from community models or LoRAs
Release AnyText-benchmark dataset and evaluation code
Release AnyWord-3M dataset and training code

Citation

@article{tuo2023anytext,
      title={AnyText: Multilingual Visual Text Generation And Editing}, 
      author={Yuxiang Tuo and Wangmeng Xiang and Jun-Yan He and Yifeng Geng and Xuansong Xie},
      year={2023},
      eprint={2311.03054},
      archivePrefix={arXiv},
      primaryClass={cs.CV}
}

Name		Name	Last commit message	Last commit date
Latest commit History 14 Commits
cldm		cldm
docs		docs
example_images		example_images
font		font
javascript		javascript
ldm		ldm
models_yaml		models_yaml
ocr_recog		ocr_recog
ocr_weights		ocr_weights
.gitignore		.gitignore
LICENSE		LICENSE
README.md		README.md
bert_tokenizer.py		bert_tokenizer.py
dataset_util.py		dataset_util.py
demo.py		demo.py
environment.yaml		environment.yaml
inference.py		inference.py
style.css		style.css
t3_dataset.py		t3_dataset.py
util.py		util.py

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

AnyText: Multilingual Visual Text Generation And Editing

📌News

💡Methodology

🛠Installation

🔮Inference

🌄Gallery

📈Evaluation

⏰TODOs

Citation

About

Releases

Packages

Languages

License

larygwil/AnyText

Folders and files

Latest commit

History

Repository files navigation

AnyText: Multilingual Visual Text Generation And Editing

📌News

💡Methodology

🛠Installation

🔮Inference

🌄Gallery

📈Evaluation

⏰TODOs

Citation

About

Resources

License

Stars

Watchers

Forks

Releases

Packages 0

Languages

Packages