Recovering editable, object-level SVG/XML code from images of structured graphics.
bash convert_figs.sh to render comparison.png.Structured images β diagrams, charts, and flowcharts β are inherently symbolic and can be compactly represented in an editable format, yet in practice they are shared as raster images and are no longer graphically editable. We present Back2Struct, which "makes structured images editable again" by directly recovering vector graphics code (SVG / XML) from image representations. Given an image of a structured graphic, Back2Struct predicts semantically object-level SVG/XML that explicitly encodes text, shapes, topology, and layout β rather than performing low-level pixel vectorization β so the code can be imported straight into tools like PowerPoint to edit, restyle, and reuse content while preserving structural fidelity. Beyond supervised fine-tuning on ground-truth SVG token sequences, we further optimize Back2Struct with reward-based learning using a composite reward that jointly encourages SVG/XML compilability, length consistency with the reference code, and structural/semantic similarity to the input. Experiments show Back2Struct improves accuracy, editability, validity, and user alignment over strong baselines.
Three contributions toward editable structured-image generation.
84,121 paired structured raster images and their SVG/XML code, plus a 1,000-sample evaluation benchmark β a large-scale, high-value corpus for editable structured-image generation.
A framework that recovers semantically meaningful, object-level vector code β explicitly modeling text, shapes, topology, and layout, not pixels.
A composite reward (syntactic validity, length fidelity, and perceptual quality) optimized with GRPO improves deployment-time validity and faithfulness.
Object-level SVG generation, warm-started by SFT and refined by GRPO.
bash convert_figs.sh to render framework.png.We cast structured-image recovery as object-level vector-code generation: instead of tracing pixels, the model emits SVG/XML primitives that name the actual shapes, text, connectors, and their arrangement. The policy is first supervised-fine-tuned on ground-truth SVG sequences, then optimized with reinforcement learning (GRPO) using a reward tailored to how the model is deployed.
The composite reward combines three complementary signals β syntactic validity (the output must compile/render), length fidelity (consistency with the reference code length), and perceptual quality (structural/semantic similarity between the rendered prediction and the target image). Group-relative advantages then steer the policy toward outputs that are simultaneously valid, concise, and visually faithful.
84K paired imageβcode examples across three sources, with a 1,000-sample held-out benchmark.
| Subset | Train (full) | Train (≤8,192 tok) | Benchmark |
|---|---|---|---|
| StarDiag-HQ | 52,875 | 52,875 | 569 |
| OpenDiag | 27,953 | 18,800 | 334 |
| NNArch | 3,293 | 1,453 | 97 |
| Total | 84,121 | 73,128 | 1,000 |
The benchmark is split by ground-truth SVG length into easy (≤2,048 tok), medium (2,048β4,096), and hard (>4,096) tiers.
bash convert_figs.sh to render dataset_overview.png.
bash convert_figs.sh to render data_samples.png.On StructHub, Back2Struct is the strongest open/trained model β and on renderable outputs it rivals frontier closed models many times its size.
| Model | Image-level | Text-level (GPTScore) | ||||||
|---|---|---|---|---|---|---|---|---|
| SR ↑ | DINO ↑ | LPIPS ↓ | SSIM ↑ | Overall ↑ | Struct ↑ | Flow ↑ | Style ↑ | |
| Overall | ||||||||
| Qwen3-VL-30B-A3B | 18.8 | 0.1566 | 0.9067 | 0.1201 | 13.29 | 0.78 | 0.68 | 0.53 |
| StarVector-8B | 31.4 | 0.2694 | 0.7936 | 0.2170 | 21.05 | 1.04 | 1.09 | 1.03 |
| Qwen2.5-VL-7B | 52.3 | 0.3527 | 0.7834 | 0.3276 | 27.56 | 1.65 | 1.32 | 1.16 |
| Qwen2.5-VL-7BSFT | 42.8 | 0.3715 | 0.7403 | 0.2891 | 33.49 | 1.80 | 1.65 | 1.57 |
| Back2Struct (RL) | 71.5 | 0.5622 | 0.6473 | 0.4966 | 51.85 | 2.87 | 2.60 | 2.31 |
| Easy | ||||||||
| Qwen3-VL-30B-A3B | 26.5 | 0.2231 | 0.8659 | 0.1702 | 19.51 | 1.15 | 1.00 | 0.78 |
| StarVector-8B | 46.1 | 0.3923 | 0.7023 | 0.3123 | 31.90 | 1.61 | 1.74 | 1.44 |
| Qwen2.5-VL-7B | 69.3 | 0.4776 | 0.7123 | 0.4196 | 39.66 | 2.29 | 2.04 | 1.61 |
| Qwen2.5-VL-7BSFT | 56.0 | 0.4866 | 0.6630 | 0.3789 | 44.58 | 2.39 | 2.30 | 1.99 |
| Back2Struct (RL) | 84.0 | 0.6384 | 0.5956 | 0.5796 | 62.63 | 3.44 | 3.25 | 2.70 |
| Medium | ||||||||
| Qwen3-VL-30B-A3B | 19.6 | 0.1631 | 0.9004 | 0.1288 | 13.38 | 0.78 | 0.68 | 0.54 |
| StarVector-8B | 37.0 | 0.3232 | 0.7533 | 0.2583 | 25.27 | 1.23 | 1.25 | 1.31 |
| Qwen2.5-VL-7B | 51.8 | 0.3546 | 0.7816 | 0.3306 | 26.88 | 1.62 | 1.25 | 1.16 |
| Qwen2.5-VL-7BSFT | 45.5 | 0.3985 | 0.7161 | 0.3091 | 36.50 | 1.98 | 1.76 | 1.74 |
| Back2Struct (RL) | 75.6 | 0.6088 | 0.6182 | 0.5275 | 55.14 | 3.08 | 2.72 | 2.47 |
| Hard | ||||||||
| Qwen3-VL-30B-A3B | 10.2 | 0.0836 | 0.9536 | 0.0614 | 7.01 | 0.41 | 0.36 | 0.28 |
| StarVector-8B | 11.1 | 0.0932 | 0.9249 | 0.0808 | 6.02 | 0.28 | 0.28 | 0.34 |
| Qwen2.5-VL-7B | 35.7 | 0.2263 | 0.8561 | 0.2331 | 16.17 | 1.03 | 0.67 | 0.72 |
| Qwen2.5-VL-7BSFT | 27.0 | 0.2297 | 0.8413 | 0.1795 | 19.43 | 1.04 | 0.89 | 0.98 |
| Back2Struct (RL) | 55.0 | 0.4399 | 0.7278 | 0.3830 | 37.83 | 2.09 | 1.83 | 1.76 |
Back2Struct is the strongest open / trained model on every tier and metric β render success 52.3%β71.5% overall β with its largest margins on the Hard tier.
| Model | DINO ↑ | LPIPS ↓ | SSIM ↑ | GPT ↑ |
|---|---|---|---|---|
| Qwen2.5-VL-7B | 0.6749 | 0.5856 | 0.6270 | 52.73 |
| Qwen3-VL-30B-A3B | 0.8347 | 0.5027 | 0.6402 | 70.88 |
| Gemini-2.5-Pro | 0.8855 | 0.4648 | 0.6532 | 77.95 |
| GPT-5 | 0.8852 | 0.4350 | 0.6573 | 77.21 |
| Back2Struct (7B) | 0.8697 | 0.3896 | 0.6738 | 78.63 |
On compiled outputs, Back2Struct achieves the best GPTScore, SSIM, and LPIPS β ahead of GPT-5 and Gemini-2.5-Pro, at a fraction of the size.
bash convert_figs.sh to render performance.png.
bash convert_figs.sh to render qualitative.png.
bash convert_figs.sh to render text2svg.png.@misc{yan2026back2structmakingstructuredimages,
title={Back2Struct: Making Structured Images Editable Again},
author={Pengyu Yan and Yixin Wu and Yunjie Tian and David Doermann},
year={2026},
eprint={2609.37016},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.37016},
}