Preprint · 2026

Back2Struct: Making Structured Images Editable Again

Recovering editable, object-level SVG/XML code from images of structured graphics.

Pengyu Yan1    Yixin Wu1    Yunjie Tian1    David Doermann1
1 University at Buffalo, SUNY
Qualitative comparison of image-to-SVG recovery across GPT-5, Qwen3-Max, Gemini-2.5-Pro, and Back2Struct
Figure not generated yet β€” run bash convert_figs.sh to render comparison.png.
Recovering the reference image (Fig. 1). A qualitative comparison across the most advanced LLMs β€” GPT-5 (Thinking), Qwen3-Max (>1T params), Gemini-2.5-Pro β€” and Back2Struct (ours). With only 7B parameters, Back2Struct reconstructs structured graphics with negligible confusion or mistakes.

Abstract

Structured images β€” diagrams, charts, and flowcharts β€” are inherently symbolic and can be compactly represented in an editable format, yet in practice they are shared as raster images and are no longer graphically editable. We present Back2Struct, which "makes structured images editable again" by directly recovering vector graphics code (SVG / XML) from image representations. Given an image of a structured graphic, Back2Struct predicts semantically object-level SVG/XML that explicitly encodes text, shapes, topology, and layout β€” rather than performing low-level pixel vectorization β€” so the code can be imported straight into tools like PowerPoint to edit, restyle, and reuse content while preserving structural fidelity. Beyond supervised fine-tuning on ground-truth SVG token sequences, we further optimize Back2Struct with reward-based learning using a composite reward that jointly encourages SVG/XML compilability, length consistency with the reference code, and structural/semantic similarity to the input. Experiments show Back2Struct improves accuracy, editability, validity, and user alignment over strong baselines.

Highlights

Three contributions toward editable structured-image generation.

84K

StructHub dataset

84,121 paired structured raster images and their SVG/XML code, plus a 1,000-sample evaluation benchmark β€” a large-scale, high-value corpus for editable structured-image generation.

SVG/XML

Object-level recovery

A framework that recovers semantically meaningful, object-level vector code β€” explicitly modeling text, shapes, topology, and layout, not pixels.

RL

Reward-based optimization

A composite reward (syntactic validity, length fidelity, and perceptual quality) optimized with GRPO improves deployment-time validity and faithfulness.

7B
parameters
84,121
image–code pairs
1,000
benchmark samples
SOTA
among open-source models

Method

Object-level SVG generation, warm-started by SFT and refined by GRPO.

Back2Struct framework overview
Figure not generated yet β€” run bash convert_figs.sh to render framework.png.
Overview of Back2Struct (Fig. 2). Given an RGB reference image (and/or text), the policy autoregressively generates object-level SVG/XML that losslessly converts into an editable document. Training is SFT warm-start followed by GRPO group-relative optimization.

We cast structured-image recovery as object-level vector-code generation: instead of tracing pixels, the model emits SVG/XML primitives that name the actual shapes, text, connectors, and their arrangement. The policy is first supervised-fine-tuned on ground-truth SVG sequences, then optimized with reinforcement learning (GRPO) using a reward tailored to how the model is deployed.

The composite reward combines three complementary signals β€” syntactic validity (the output must compile/render), length fidelity (consistency with the reference code length), and perceptual quality (structural/semantic similarity between the rendered prediction and the target image). Group-relative advantages then steer the policy toward outputs that are simultaneously valid, concise, and visually faithful.

StructHub Dataset

84K paired image–code examples across three sources, with a 1,000-sample held-out benchmark.

StructHub statistics (Table 1). Training uses examples with ≤8,192 tokens; the benchmark is held out.
SubsetTrain (full)Train (≤8,192 tok)Benchmark
StarDiag-HQ52,87552,875569
OpenDiag27,95318,800334
NNArch3,2931,45397
Total84,12173,1281,000

The benchmark is split by ground-truth SVG length into easy (≤2,048 tok), medium (2,048–4,096), and hard (>4,096) tiers.

StructHub dataset statistics
Figure not generated yet β€” run bash convert_figs.sh to render dataset_overview.png.
Dataset statistics (Fig. 3). Token-length distributions, content profiles, and difficulty compositions of StarDiag-HQ, OpenDiag, and NNArch.
Examples from the StructHub dataset
Figure not generated yet β€” run bash convert_figs.sh to render data_samples.png.
Dataset examples (Fig. 5). Each raster image is paired with an editable SVG counterpart, providing high-quality pixel–code supervision.

Results

On StructHub, Back2Struct is the strongest open/trained model β€” and on renderable outputs it rivals frontier closed models many times its size.

Image-level and text-level evaluation on StructHub across difficulty tiers (Table 2, render-gated: a non-rendering prediction scores 0 on DINO/SSIM/GPT and 1 on LPIPS). SR is render success (%); GPTScore Overall is 0–100, each sub-axis 0–5. Best and second-best per column; our model is shaded.
Model Image-level Text-level (GPTScore)
SR ↑DINO ↑LPIPS ↓SSIM ↑ Overall ↑Struct ↑Flow ↑Style ↑
Overall
Qwen3-VL-30B-A3B18.80.15660.90670.120113.290.780.680.53
StarVector-8B31.40.26940.79360.217021.051.041.091.03
Qwen2.5-VL-7B52.30.35270.78340.327627.561.651.321.16
Qwen2.5-VL-7BSFT42.80.37150.74030.289133.491.801.651.57
Back2Struct (RL)71.50.56220.64730.496651.852.872.602.31
Easy
Qwen3-VL-30B-A3B26.50.22310.86590.170219.511.151.000.78
StarVector-8B46.10.39230.70230.312331.901.611.741.44
Qwen2.5-VL-7B69.30.47760.71230.419639.662.292.041.61
Qwen2.5-VL-7BSFT56.00.48660.66300.378944.582.392.301.99
Back2Struct (RL)84.00.63840.59560.579662.633.443.252.70
Medium
Qwen3-VL-30B-A3B19.60.16310.90040.128813.380.780.680.54
StarVector-8B37.00.32320.75330.258325.271.231.251.31
Qwen2.5-VL-7B51.80.35460.78160.330626.881.621.251.16
Qwen2.5-VL-7BSFT45.50.39850.71610.309136.501.981.761.74
Back2Struct (RL)75.60.60880.61820.527555.143.082.722.47
Hard
Qwen3-VL-30B-A3B10.20.08360.95360.06147.010.410.360.28
StarVector-8B11.10.09320.92490.08086.020.280.280.34
Qwen2.5-VL-7B35.70.22630.85610.233116.171.030.670.72
Qwen2.5-VL-7BSFT27.00.22970.84130.179519.431.040.890.98
Back2Struct (RL)55.00.43990.72780.383037.832.091.831.76

Back2Struct is the strongest open / trained model on every tier and metric β€” render success 52.3%β†’71.5% overall β€” with its largest margins on the Hard tier.

Comparison on successfully compiled outputs (Table 3). On renderable SVGs, the 7B Back2Struct matches or beats frontier closed models. Best / second-best.
ModelDINO ↑LPIPS ↓SSIM ↑GPT ↑
Qwen2.5-VL-7B0.67490.58560.627052.73
Qwen3-VL-30B-A3B0.83470.50270.640270.88
Gemini-2.5-Pro0.88550.46480.653277.95
GPT-50.88520.43500.657377.21
Back2Struct (7B)0.86970.38960.673878.63

On compiled outputs, Back2Struct achieves the best GPTScore, SSIM, and LPIPS β€” ahead of GPT-5 and Gemini-2.5-Pro, at a fraction of the size.

Image-level performance comparison across dataset sources
Figure not generated yet β€” run bash convert_figs.sh to render performance.png.
Performance by source (Fig. 4). Image-level performance comparison across the three StructHub sources (StarDiag-HQ, OpenDiag, NNArch).

Qualitative Samples

Qualitative results of Back2Struct
Figure not generated yet β€” run bash convert_figs.sh to render qualitative.png.
Image-to-SVG results (Fig. 10). Gray denotes ground-truth SVGs; green indicates Back2Struct predictions across schemas, architecture diagrams, workflows, plots, and hierarchies.
Text-to-SVG comparison
Figure not generated yet β€” run bash convert_figs.sh to render text2svg.png.
Text-to-SVG (Fig. 11). Given the same prompt, the 7B Back2Struct produces diagrams closer to the ground truth than much larger LLMs β€” the reward-driven fine-tuning transfers from image- to text-conditioned generation.

BibTeX

@misc{yan2026back2structmakingstructuredimages,
      title={Back2Struct: Making Structured Images Editable Again},
      author={Pengyu Yan and Yixin Wu and Yunjie Tian and David Doermann},
      year={2026},
      eprint={2609.37016},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.37016},
}