Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ
Review for "Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ"
Summary
This paper introduces Edit2TikZ, a comprehensive benchmark for evaluating instruction-guided scientific figure editing using TikZ code. The work addresses a critical gap in existing benchmarks, which often focus on reconstruction or generation rather than editing. The proposed benchmark includes real-world and synthetic examples, supports both text and visual localization, and provides detailed annotations for multi-step edits. The authors also present an evaluation framework that measures both edit correctness and preservation of non-target content.
Mathematical/empirical assessment
The paper presents a well-structured approach to scientific figure editing, with clear definitions of atomic edit operations and a thorough data construction process. The evaluation metrics—particularly the human-aligned RestorationScore (RS) and EditCorrectnessScore (ECS)—offer a nuanced way to assess model performance beyond simple compilation success. The experiments show that even state-of-the-art models struggle with this task, highlighting the difficulty of combining visual understanding, instruction following, and code generation.
Strengths
The paper makes a strong case for the importance of scientific figure editing as a task that goes beyond simple reconstruction or generation. The benchmark is diverse, well-documented, and includes both real-world and synthetic examples. The evaluation framework is thoughtful and aligns closely with human judgment, which is crucial for tasks involving complex visual and semantic changes. The training strategy with TikZEditMix and curriculum learning demonstrates practical improvements, especially for compact models.
Concerns
While the paper is well-executed, some aspects could be clarified. For instance, the exact nature of the "multi-step editing" and how the step-level annotations are generated is not fully detailed. Additionally, while the human alignment study is promising, it would be helpful to see more information about the annotation process and inter-rater reliability beyond what is provided. The paper also does not discuss potential limitations of the evaluation framework, such as how it handles ambiguous or subjective edits.
Final decision
Strong accept
The paper makes a valuable contribution to the field by introducing a robust and challenging benchmark for scientific figure editing. The methodology is sound, the evaluation is thorough, and the results provide important insights into the current capabilities and limitations of MLLMs in this domain. The work is well-suited for publication and will likely serve as a useful reference for future research.