OneFont: A Unified Agent for End-to-End Font Creation
From Manual Model Selection to OneFont

Motivation
Existing diffusion and MLLM-based font generation systems often focus on a single editing task. Supporting a new script or generation workflow can require substantial task-specific fine-tuning, while selecting the right model and repairing an imperfect result still depend on expert intervention.
OneFont treats end-to-end font creation as an MLLM agent task. Given a user request, it selects suitable models and tools, generates candidate results, verifies readability and style, and performs local refinement when needed.
Method Overview
OneFont converts each request into a reasoning trace and structured tool calls. Its generation library covers text-to-art fonts, style transfer, vector font generation, retrieval, handwriting generation, and local editing. Supervised fine-tuning teaches the agent to plan and invoke these tools, while GRPO-based preference alignment improves the quality of its generation strategy. During inference, a graph-based planner verifies intermediate results, backtracks when necessary, and applies local refinement.
Method Breakdown
The OneFont framework combines a multimodal generation tool library, two-stage SFT and preference-alignment training, and a graph-based planner for verification and backtracking.

Qualitative Results
Across text-rich generation tasks, OneFont produces readable text while matching the requested object, composition, and visual style. The comparison includes general image generators and dedicated text rendering models.
