OneFont: A Unified Agent for End-to-End Font Creation

Mar 2026·
Yingxin Lai
Yufei Liu
Yufei Liu
,
Guoqing Yang
,
Jiaxing Chai
,
Zhiming Luo
,
Shaozi Li
· 2 min read
Type
Publication
In Proceedings of the AAAI Conference on Artificial Intelligence, 40(1), 552-560

From Manual Model Selection to OneFont

Comparison of a default font model, expert model selection, and OneFont for generating a Planet Earth design
OneFont completes model selection and refinement in a single agent workflow.

Motivation

Existing diffusion and MLLM-based font generation systems often focus on a single editing task. Supporting a new script or generation workflow can require substantial task-specific fine-tuning, while selecting the right model and repairing an imperfect result still depend on expert intervention.

OneFont treats end-to-end font creation as an MLLM agent task. Given a user request, it selects suitable models and tools, generates candidate results, verifies readability and style, and performs local refinement when needed.

Method Overview

OneFont converts each request into a reasoning trace and structured tool calls. Its generation library covers text-to-art fonts, style transfer, vector font generation, retrieval, handwriting generation, and local editing. Supervised fine-tuning teaches the agent to plan and invoke these tools, while GRPO-based preference alignment improves the quality of its generation strategy. During inference, a graph-based planner verifies intermediate results, backtracks when necessary, and applies local refinement.

Method Breakdown

The OneFont framework combines a multimodal generation tool library, two-stage SFT and preference-alignment training, and a graph-based planner for verification and backtracking.

OneFont training framework, generation tool library, and graph-based planner
Training framework and graph-based inference planner
01
Tool LibraryText, image, and answer tools for multimodal font creation
02
Model TrainingSFT builds tool-use ability, then GRPO aligns generation preferences
03
Model InferenceThe planner verifies results, backtracks when needed, and refines the output

Qualitative Results

Across text-rich generation tasks, OneFont produces readable text while matching the requested object, composition, and visual style. The comparison includes general image generators and dedicated text rendering models.

Qualitative comparison of OneFont with AnyText, DALL-E 3, Ideogram, PixArt-alpha, SDXL, and TextDiffuser-2
Qualitative comparison on cakes, T-shirts, book covers, and character images with embedded text. The final column shows OneFont results.