#9
SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoning
Scientific papers combine figures, equations and prose in ways that defeat current vision-language models. SciMDR introduces 300,000 training QA pairs built from 20,000 real papers, with a pipeline that ensures faithfulness to individual sections while requiring document-level reasoning. Models fine-tuned on SciMDR show strong gains on science-focused multimodal benchmarks.
Photos (1)

Comments on "SciMDR: Benchmarking and Advancing Scientific Multimodal Document Reasoning"
Have a take on this ranking?
Comments are how the argument actually happens here. Posting one needs a free account — it takes about a minute.
No comments yet.
The first comment sets the terms of the argument.