Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ
I see where you are coming from, but I think the answer is more mixed.
The paper presents a well-structured and timely contribution to the field of scientific figure editing, addressing a critical gap in existing benchmarks that often focus on reconstruction or generation rather than instruction-guided editing. The introduction of Edit2TikZ with its diverse dataset, multi-step editing capabilities, and human-aligned evaluation metrics is a significant step forward. The training strategy with TikZEditMix and curriculum learning demonstrates practical improvements, particularly for compact models.
That said, the paper could benefit from more in-depth analysis of why certain edit types are more challenging than others and how different model architectures affect performance. While the results are promising, a deeper exploration of these factors would strengthen the paper's impact. The comparison with other benchmarks is limited, and further discussion of trade-offs between model size and performance would be valuable.
The part I find convincing is the thorough evaluation framework, which provides a nuanced way to assess model performance beyond simple compilation success. The human-aligned metrics, RS and ECS, are particularly insightful and align closely with human judgment, which is crucial for tasks involving complex visual and semantic changes.
I have one genuine question: Could the authors provide more details on how the step-level annotations are generated and validated? This would help in understanding the reliability and consistency of the dataset.
Strong accept