{"id":9453,"date":"2025-12-05T22:09:33","date_gmt":"2025-12-05T22:09:33","guid":{"rendered":"https:\/\/techtrendfeed.com\/?p=9453"},"modified":"2025-12-05T22:09:33","modified_gmt":"2025-12-05T22:09:33","slug":"so-bench-a-structural-output-analysis-of-multimodal-llms","status":"publish","type":"post","link":"https:\/\/techtrendfeed.com\/?p=9453","title":{"rendered":"SO-Bench: A Structural Output Analysis of Multimodal LLMs"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p>Multimodal massive language fashions (MLLMs) are more and more deployed in real-world, agentic settings the place outputs should not solely be right, but in addition conform to predefined information schemas. Regardless of latest progress in structured era in textual area, there&#8217;s nonetheless no benchmark that systematically evaluates schema-grounded info extraction and reasoning over visible inputs. On this work, we conduct a complete examine of visible structural output capabilities for MLLMs with our rigorously designed SO-Bench benchmark. Overlaying 4 visible domains, together with UI screens, pure photographs, paperwork, and charts, SO-Bench is constructed from over 6.5K various JSON schemas and 1.8K curated image-schema pairs with human-verified high quality. Benchmarking experiments on open-sourced and frontier proprietary fashions reveal persistent gaps in predicting correct, schema compliant outputs, highlighting the necessity for higher multimodal structured reasoning. Past benchmarking, we additional conduct coaching experiments to largely enhance the mannequin\u2019s structured output functionality. We plan to make the benchmark obtainable to the neighborhood.<\/p>\n<figure id=\"figure1\" class=\"\" aria-label=\"Figure 1\">\n<div class=\"bg-gray-light text-base rounded\"><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/mlr.cdn-apple.com\/media\/introduction_converted_2da889b02d.png\" aria-label=\"Composite figure showing two panels: left, the SO-Bench multi-stage data generation pipeline including schema, intent, and response stages with model and human expert involvement; right, benchmarking results comparing open-source and proprietary frontier models.\" tabindex=\"-1\" target=\"_blank\" class=\"mt-0\"><img decoding=\"async\" src=\"https:\/\/mlr.cdn-apple.com\/media\/introduction_converted_2da889b02d.png\" alt=\"Composite figure showing two panels: left, the SO-Bench multi-stage data generation pipeline including schema, intent, and response stages with model and human expert involvement; right, benchmarking results comparing open-source and proprietary frontier models.\" loading=\"lazy\" class=\"bg-gray-light\"\/><\/a><\/div><figcaption class=\"muted\" aria-hidden=\"true\">Determine 1: Left: Overview of the multi-stage information era pipeline for SO-Bench, together with schema era, consumer intent era, and response era levels. At every stage, proprietary frontier fashions comparable to GPT-5 and Gemini-2.5-Professional act as turbines with rigorously designed prompts. Human area consultants assessment information from every stage earlier than it progresses to the following. Previous to schema era, enter photographs and JSON schemas are embedded utilizing a CLIP mannequin for embedding search. Proper: Benchmarking outcomes amongst a number of open-source fashions and proprietary frontier fashions.<\/figcaption><\/figure>\n<figure id=\"figure2\" class=\"\" aria-label=\"Figure 2\">\n<div class=\"bg-gray-light text-base rounded\"><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/mlr.cdn-apple.com\/media\/data_generation_pipeline_converted_6ed8a55c36.png\" aria-label=\"Diagram of the SO-Bench data generation pipeline showing schema generation, user intent generation, response generation, and CLIP-based embedding search with human expert checks at each stage.\" tabindex=\"-1\" target=\"_blank\" class=\"mt-0\"><img decoding=\"async\" src=\"https:\/\/mlr.cdn-apple.com\/media\/data_generation_pipeline_converted_6ed8a55c36.png\" alt=\"Diagram of the SO-Bench data generation pipeline showing schema generation, user intent generation, response generation, and CLIP-based embedding search with human expert checks at each stage.\" loading=\"lazy\" class=\"bg-gray-light\"\/><\/a><\/div><figcaption class=\"muted\" aria-hidden=\"true\">Determine 2: Overview of the multi-stage information era pipeline for SO-Bench, together with schema era, consumer intent era, and response era levels. At every stage, proprietary frontier fashions comparable to GPT-5 and Gemini-2.5-Professional act as turbines with rigorously designed prompts. Human area consultants assessment information from every stage earlier than it progresses to the following. Previous to schema era, enter photographs and JSON schemas are embedded utilizing a CLIP mannequin for embedding search.<\/figcaption><\/figure>\n<\/div>\n\n","protected":false},"excerpt":{"rendered":"<p>Multimodal massive language fashions (MLLMs) are more and more deployed in real-world, agentic settings the place outputs should not solely be right, but in addition conform to predefined information schemas. Regardless of latest progress in structured era in textual area, there&#8217;s nonetheless no benchmark that systematically evaluates schema-grounded info extraction and reasoning over visible inputs. [&hellip;]<\/p>\n","protected":false},"author":2,"featured_media":9455,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[55],"tags":[608,1112,306,5200,6772,6773],"class_list":["post-9453","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-machine-learning","tag-evaluation","tag-llms","tag-multimodal","tag-output","tag-sobench","tag-structural"],"_links":{"self":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts\/9453","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=9453"}],"version-history":[{"count":1,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts\/9453\/revisions"}],"predecessor-version":[{"id":9454,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts\/9453\/revisions\/9454"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/media\/9455"}],"wp:attachment":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=9453"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=9453"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=9453"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}<!-- This website is optimized by Airlift. Learn more: https://airlift.net. Template:. Learn more: https://airlift.net. Template: 69d9690a190636c2e0989534. Config Timestamp: 2026-04-10 21:18:02 UTC, Cached Timestamp: 2026-08-02 12:37:23 UTC -->