Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ
Summary
This paper introduces Edit2TikZ, a comprehensive benchmark for evaluating instruction-guided scientific figure editing using TikZ code. The dataset includes 1,548 diverse samples with multi-step edits, visual and textual localization, and human-aligned evaluation metrics. The authors also present a training strategy that improves performance on compact models like Qwen3.5-4B.
Mathematical/empirical assessment
The paper provides clear definitions of edit operations and a well-structured evaluation framework. The results show significant improvements in compilation success and edit correctness after applying the proposed training strategy. However, the paper lacks detailed analysis of why certain edit types are more challenging than others, and it doesn't explore how different model architectures affect performance.
Strengths
The benchmark is well-designed, covering a wide range of edit types and including both real-world and synthetic data. The human-aligned evaluation metrics are a strong point, as they address limitations of existing automated metrics. The training strategy shows promising results, especially for compact models.
Concerns
The paper could benefit from more in-depth analysis of the difficulty of different edit operations and their impact on model performance. Additionally, the comparison with other benchmarks is limited, and the paper doesn't fully explore the trade-offs between model size and performance.
Final decision
Strong accept