<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0"><channel><title>Benchmark Radar updates</title><link>https://benchmark-radar.com/</link><description>New and updated public AI benchmarks indexed by Benchmark Radar.</description><language>en</language><lastBuildDate>Sun, 30 Aug 2026 00:00:00 +0000</lastBuildDate><atom:link xmlns:atom="http://www.w3.org/2005/Atom" href="https://benchmark-radar.com/feed.xml" rel="self" type="application/rss+xml"/><item><title>WhatIfBench</title><link>https://benchmark-radar.com/benchmarks/bm_the-illusion-of-textit-what-if-evaluating-_2b6da5c9/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_the-illusion-of-textit-what-if-evaluating-_2b6da5c9/</guid><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><description>A diagnostic benchmark of 220 open-domain what-if questions across STEM, HSS, and Hybrid scenarios, evaluated with PRISM metrics on causal graphs and explanatory adequacy.</description></item><item><title>VisTarget-Bench</title><link>https://benchmark-radar.com/benchmarks/bm_vistarget-bench_aaa3c19c/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_vistarget-bench_aaa3c19c/</guid><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><description>A 150-task human-verified benchmark pairing questions with held-out target images to separate image-retrieval failures from visual-perception failures in multimodal search agents.</description></item><item><title>SPAR-Bench</title><link>https://benchmark-radar.com/benchmarks/bm_spar-bench_b75502fb/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_spar-bench_b75502fb/</guid><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><description>Evaluates spatial reasoning in medical vision encoders via eight probes over multi-organ abdominal CT covering coordinate localization, relational reasoning, and spatial queries.</description></item><item><title>PCBnet</title><link>https://benchmark-radar.com/benchmarks/bm_pcbnet_ffc56858/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_pcbnet_ffc56858/</guid><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><description>Printed circuit boards (PCBs) are fundamental to modern electronic systems, yet AI-driven PCB design automation remains constrained by the lack of large-scale paired schematic-netlist datasets.</description></item><item><title>NumBench</title><link>https://benchmark-radar.com/benchmarks/bm_numbench_cf80a48e/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_numbench_cf80a48e/</guid><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><description>Text-to-image (T2I) models often generate the wrong number of objects, yet existing benchmarks are too small or weakly controlled to explain why.</description></item><item><title>MOSAIC</title><link>https://benchmark-radar.com/benchmarks/bm_beyond-global-scalars-synergizing-token-le_0d7c6f3e/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_beyond-global-scalars-synergizing-token-le_0d7c6f3e/</guid><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><description>Adversarial AIGC text detection benchmark with 16,000 samples spanning a full-granularity attack spectrum.</description></item><item><title>MD-VQA</title><link>https://benchmark-radar.com/benchmarks/bm_post-training-vlms-for-video-mistake-detec_58e7640e/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_post-training-vlms-for-video-mistake-detec_58e7640e/</guid><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><description>Tests video models on detecting whether a step was executed correctly according to its description, for seen and unseen actions.</description></item><item><title>LoopArena</title><link>https://benchmark-radar.com/benchmarks/bm_looparena_97bc3500/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_looparena_97bc3500/</guid><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><description>Loop Engineering is emerging as a practice for organizing development work around coding agents.</description></item><item><title>GenIaC-SecBench</title><link>https://benchmark-radar.com/benchmarks/bm_geniac-secbench_393dfc82/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_geniac-secbench_393dfc82/</guid><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><description>Large language models are increasingly used to author Infrastructure-as-Code (IaC), where a single insecure default can be deployed directly into production.</description></item><item><title>FinExam-10K</title><link>https://benchmark-radar.com/benchmarks/bm_finexam-10k-when-retrieval-helps-financial_fa6458fb/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_finexam-10k-when-retrieval-helps-financial_fa6458fb/</guid><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><description>Evaluates financial reasoning across CFA Levels I-III and FRM Parts I-II with 10,198 expert-reannotated questions, releasing 5,110 public items and maintaining a leaderboard on 5,088 held-out items.</description></item><item><title>EvoHarmBench</title><link>https://benchmark-radar.com/benchmarks/bm_evoharmbench-breaking-content-moderation-w_fad69c38/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_evoharmbench-breaking-content-moderation-w_fad69c38/</guid><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><description>Dynamic adversarial evaluation framework that evolves evasion strategies across 229 semantic sub-clusters from 5,002 real-world samples to assess content moderation systems.</description></item><item><title>ElephantBench</title><link>https://benchmark-radar.com/benchmarks/bm_elephantbench_79ce069e/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_elephantbench_79ce069e/</guid><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><description>Closed-book knowledge probe with 1,094 questions evaluating recall of multiple divergent accounts for long-tail facts, with fixed C/P/F/K scoring metrics.</description></item><item><title>Copper Tube Defect Dataset</title><link>https://benchmark-radar.com/benchmarks/bm_cf-yolo-context-aware-feature-refinement-f_82f82c68/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_cf-yolo-context-aware-feature-refinement-f_82f82c68/</guid><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><description>Evaluates object detection of micro-defects on copper tube surfaces using 1,847 images and 4,898 bounding box instances across defect types in industrial inspection.</description></item><item><title>CoCoBench</title><link>https://benchmark-radar.com/benchmarks/bm_cocobench_0d6f9f05/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_cocobench_0d6f9f05/</guid><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><description>Agent systems powered by multimodal large language models (MLLMs) have advanced rapidly in recent years, yet existing embodied-agent benchmarks still lack fine-grained diagnostics for multi-agent coordination.</description></item><item><title>CNeo-Bench</title><link>https://benchmark-radar.com/benchmarks/bm_cneo-bench_bbdf3de1/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_cneo-bench_bbdf3de1/</guid><pubDate>Fri, 28 Aug 2026 00:00:00 +0000</pubDate><description>Chinese neologisms exploit diverse and unique linguistic mechanisms, such as phonetic substitution (e.g., 886 for ``bye-bye'') and visual character decomposition that are rare in other languages.</description></item><item><title>RuleWeaver</title><link>https://benchmark-radar.com/benchmarks/bm_ruleweaver-benchmarking-rule-centered-scen_2379b2dc/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_ruleweaver-benchmarking-rule-centered-scen_2379b2dc/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>Evaluates rule-centered scenario reasoning through corpus-derived IF-THEN rules, composed into QA instances with rubric-based answer quality, rule recall, and rule precision scoring.</description></item><item><title>ReViCo</title><link>https://benchmark-radar.com/benchmarks/bm_revico_34e02af8/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_revico_34e02af8/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>Vision Language Models (VLMs) have shown great success in general visual tasks, yet they still struggle to deeply understand text within images.</description></item><item><title>RATIO</title><link>https://benchmark-radar.com/benchmarks/bm_ratio-a-benchmark-for-retrieval-across-typ_4be3618a/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_ratio-a-benchmark-for-retrieval-across-typ_4be3618a/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>Evaluates retrieval models on three ideation moves—Address, Broaden, Specify—using relevance judgments derived from full-text scientific papers across computer science literature.</description></item><item><title>R2M-Bench</title><link>https://benchmark-radar.com/benchmarks/bm_r2m-bench_1ea8e143/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_r2m-bench_1ea8e143/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>High similarity between first-visit and return frames does not necessarily show that a video world model remembered the scene; the intervening rollout may simply have changed very little.</description></item><item><title>PLCBENCH</title><link>https://benchmark-radar.com/benchmarks/bm_plcbench_fb7005e5/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_plcbench_fb7005e5/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>A hardware-in-the-loop framework evaluating LLM agents across commercial PLCs and physical process simulations, with deterministic scoring for PLC interaction, process-linked manipulation, and sustained physical impact.</description></item><item><title>PAWBench</title><link>https://benchmark-radar.com/benchmarks/bm_pawbench_0877eb45/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_pawbench_0877eb45/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>Evaluates video generators as stochastic samplers of world dynamics across 50 scenarios, comparing the distribution of possible behaviors under identical initial observations and actions.</description></item><item><title>Multi2AV-Safety</title><link>https://benchmark-radar.com/benchmarks/bm_multi2av-safety_13dc6c17/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_multi2av-safety_13dc6c17/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>Audio-video generation is rapidly moving from prompt-driven synthesis toward multimodal conditioning, where text, images, audio, and video can jointly shape the generated output.</description></item><item><title>MCR-Bench</title><link>https://benchmark-radar.com/benchmarks/bm_mcr-bench_7a6da290/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_mcr-bench_7a6da290/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>In real-world software development, code review typically involves iterative interactions between developers and reviewers to improve software quality, making the process costly and time-consuming.</description></item><item><title>HUG-VIS</title><link>https://benchmark-radar.com/benchmarks/bm_hug-vis_24e7fc8b/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_hug-vis_24e7fc8b/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>Visual intelligence seeks to perceive, interpret, and synthesize the visual world and is central to modern computer vision.</description></item><item><title>FaulT-Bench</title><link>https://benchmark-radar.com/benchmarks/bm_fault-bench_59abb402/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_fault-bench_59abb402/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>A benchmark of 200 troubleshooting scenarios across eight network topologies evaluating network troubleshooting LLM agents under unreliable user tickets, including false fault reports and incorrect device attribution.</description></item><item><title>ESRP-Bench</title><link>https://benchmark-radar.com/benchmarks/bm_esrp-bench_4847c977/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_esrp-bench_4847c977/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>This paper introduces Embodied Scene Rearrangement Planning (ESRP), a novel task requiring embodied agents to rearrange furniture in 3D scenes to match a target configuration using only egocentric observations and a top-down target layout.</description></item><item><title>DuMateBench</title><link>https://benchmark-radar.com/benchmarks/bm_dumatebench_7a085138/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_dumatebench_7a085138/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>Autonomous agents are increasingly adopted to complete complex, multi-tool workflows in real-world settings.</description></item><item><title>DEEPCHART</title><link>https://benchmark-radar.com/benchmarks/bm_deepchart_94c0e719/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_deepchart_94c0e719/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>Faithful chart generation in real-world data-science workflows requires grounding visualizations in scattered evidence, computing chart-ready quantities, and rendering them accurately.</description></item><item><title>CorporateBench</title><link>https://benchmark-radar.com/benchmarks/bm_corporatebench_6a07b031/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_corporatebench_6a07b031/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>LLMs are increasingly able to answer complex questions about enterprise-scale document collections.</description></item><item><title>BrailleBench</title><link>https://benchmark-radar.com/benchmarks/bm_braillebench_3cc849b8/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_braillebench_3cc849b8/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>Although Large language models (LLMs) mediate access to knowledge and computational assistance, their capabilities should benefit vulnerable groups in the same way.</description></item><item><title>BekchiAI-Benchmark</title><link>https://benchmark-radar.com/benchmarks/bm_bekchiai-measuring-observing-and-controlli_e45a38cc/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_bekchiai-measuring-observing-and-controlli_e45a38cc/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>BekchiAI-Benchmark evaluates LLM agent skills using 2,057 deterministic tool-using ReAct tasks across 7 categories, scoring accuracy plus tool-call adherence, URL hallucination, source-match, and token cost.</description></item><item><title>Behavior2Trip</title><link>https://benchmark-radar.com/benchmarks/bm_behavior2trip_b70cb42e/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_behavior2trip_b70cb42e/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>Travel planning agents assist users in generating personalized travel plans by modeling their individual preferences.</description></item><item><title>BTS-AgentBench</title><link>https://benchmark-radar.com/benchmarks/bm_bts-agentbench_70852571/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_bts-agentbench_70852571/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>Evaluates multi-turn agent performance on 532 telemetry-derived tasks across train/dev/test splits, with additional XAI4HEAT episodes, using deterministic replayable construction and verifiable gold answers.</description></item><item><title>BALMS</title><link>https://benchmark-radar.com/benchmarks/bm_balms_57758cdd/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_balms_57758cdd/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>Mental health assessment relies on episodic self-report scales, which convert subjective states such as stress into numerical scores but provide only sparse snapshots of wellbeing.</description></item><item><title>Ancient-Bench</title><link>https://benchmark-radar.com/benchmarks/bm_ancient-bench_967fba8e/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_ancient-bench_967fba8e/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>Ancient Chinese artifact text recognition is fundamental to heritage digitization, and benchmarks for ancient texts are essential for evaluating current model capabilities.</description></item><item><title>AgentJudgeBench</title><link>https://benchmark-radar.com/benchmarks/bm_agentjudgebench_8ae8da3e/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_agentjudgebench_8ae8da3e/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined.</description></item><item><title>4DSynth-Nav</title><link>https://benchmark-radar.com/benchmarks/bm_4dsynth-controllable-procedural-world-synt_fa98bfb4/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_4dsynth-controllable-procedural-world-synt_fa98bfb4/</guid><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><description>Evaluates embodied agents on interactive navigation tasks in procedurally generated 4D environments with independently tunable difficulty axes.</description></item><item><title>XREPOTEST</title><link>https://benchmark-radar.com/benchmarks/bm_xrepotest_04d9b86c/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_xrepotest_04d9b86c/</guid><pubDate>Wed, 26 Aug 2026 00:00:00 +0000</pubDate><description>Evaluates multilingual repository-level unit test generation across Rust, Go, Julia, PHP, and Ruby using containerized execution and context augmentation strategies.</description></item><item><title>Video-IFBench</title><link>https://benchmark-radar.com/benchmarks/bm_video-ifbench_4fc7a0ab/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_video-ifbench_4fc7a0ab/</guid><pubDate>Wed, 26 Aug 2026 00:00:00 +0000</pubDate><description>Evaluates instruction following of multimodal LLMs in video understanding with 1.5K samples across constraint categories.</description></item><item><title>VGA-BenchV2</title><link>https://benchmark-radar.com/benchmarks/bm_vga-benchv2_d6dfef90/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_vga-benchv2_d6dfef90/</guid><pubDate>Wed, 26 Aug 2026 00:00:00 +0000</pubDate><description>Evaluates video generation quality and aesthetic value with 1,016 prompts, 60,000 videos, and 36,000 task-level annotations.</description></item><item><title>VBVR-Pro</title><link>https://benchmark-radar.com/benchmarks/bm_vbvr-pro_41ab73db/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_vbvr-pro_41ab73db/</guid><pubDate>Wed, 26 Aug 2026 00:00:00 +0000</pubDate><description>Provides 300 procedurally generated tasks for native visual reasoning through generation, with verifiable reward scorers and controlled modality comparisons.</description></item><item><title>SciMIF</title><link>https://benchmark-radar.com/benchmarks/bm_scimif_1f2ceaa4/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_scimif_1f2ceaa4/</guid><pubDate>Wed, 26 Aug 2026 00:00:00 +0000</pubDate><description>Evaluates instruction following of MLLMs across five scientific disciplines with a taxonomy of 10 constraint groups.</description></item><item><title>SCALE-QA</title><link>https://benchmark-radar.com/benchmarks/bm_reconstructing-the-right-episode-evaluatin_98f96c99/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_reconstructing-the-right-episode-evaluatin_98f96c99/</guid><pubDate>Wed, 26 Aug 2026 00:00:00 +0000</pubDate><description>SCALE-QA evaluates interleaved conversational memory using 3,000 audited multiple-choice questions across 10 domains, where correct answers depend on causally related evidence from earlier turns in flat unsegmented threads.</description></item><item><title>PIVOT</title><link>https://benchmark-radar.com/benchmarks/bm_pivot-a-multi-trajectory-dataset-and-testb_3fccb501/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_pivot-a-multi-trajectory-dataset-and-testb_3fccb501/</guid><pubDate>Wed, 26 Aug 2026 00:00:00 +0000</pubDate><description>Evaluates novel-view synthesis under diverse camera trajectories, measured vs optimized poses, and calibrated vs optimized intrinsics using five real-world scenes.</description></item><item><title>OmniPhys</title><link>https://benchmark-radar.com/benchmarks/bm_omniphys_ae437965/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_omniphys_ae437965/</guid><pubDate>Wed, 26 Aug 2026 00:00:00 +0000</pubDate><description>Evaluates multimodal physics understanding, reasoning, and generation on 15,246 questions with 19,850 images from middle-school to university levels.</description></item><item><title>MathAdv</title><link>https://benchmark-radar.com/benchmarks/bm_mathadv_85c33564/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_mathadv_85c33564/</guid><pubDate>Wed, 26 Aug 2026 00:00:00 +0000</pubDate><description>Evaluates theorem proving and auxiliary tasks including multiple-choice, fill-in-the-blank, and reformulation robustness across 13 mathematical domains.</description></item><item><title>MMJailBench</title><link>https://benchmark-radar.com/benchmarks/bm_mmjailbench_1b3e3a1a/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_mmjailbench_1b3e3a1a/</guid><pubDate>Wed, 26 Aug 2026 00:00:00 +0000</pubDate><description>Evaluates multimodal jailbreak vulnerabilities by factorizing harmful intent, prompt framing, visual semantics, and instruction carrier.</description></item><item><title>FinRiskAtlas</title><link>https://benchmark-radar.com/benchmarks/bm_finriskatlas_4424cb99/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_finriskatlas_4424cb99/</guid><pubDate>Wed, 26 Aug 2026 00:00:00 +0000</pubDate><description>Evaluates Chinese financial LLMs on static operation execution across 9,742 instances and evidence-state control via FinRisk-Ask using pre-action trajectory states.</description></item><item><title>FedCMAPSS</title><link>https://benchmark-radar.com/benchmarks/bm_fedcmapss_9057f5e4/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_fedcmapss_9057f5e4/</guid><pubDate>Wed, 26 Aug 2026 00:00:00 +0000</pubDate><description>Data-driven prognostics and health management has emerged as a key enabler for Industry 4.0, yet the development of robust remaining useful life (RUL) estimation models is often limited by the scarcity of run-to-failure data.</description></item><item><title>EASEL</title><link>https://benchmark-radar.com/benchmarks/bm_paint-what-you-see-benchmarking-dexterous-_f2ccb455/</link><guid isPermaLink="true">https://benchmark-radar.com/benchmarks/bm_paint-what-you-see-benchmarking-dexterous-_f2ccb455/</guid><pubDate>Wed, 26 Aug 2026 00:00:00 +0000</pubDate><description>Evaluates dexterous visual tool use through reference-guided painting, semantic annotation, handwriting, and path planning tasks.</description></item></channel></rss>
