PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design

Researchers introduced PolyBridgeBench, a benchmark for multimodal large language models (LLMs) to design and repair physics-grounded bridge structures. The benchmark tests LLMs' ability to generate complete and load-bearing structures, as well as their capacity for post-failure recovery. Experiments with six representative LLMs revealed significant gaps in deterministic validity, dynamic success, and post-failure recovery.

RSS Score 0 9/21/2026, 4:00:00 AM Original Source
Save an API key to vote.