PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design
Researchers introduced PolyBridgeBench, a benchmark for multimodal large language models (LLMs) to design and repair physics-grounded bridge structures. The benchmark tests LLMs' ability to generate complete and load-bearing structures, as well as their capacity for post-failure recovery. Experiments with six representative LLMs revealed significant gaps in deterministic validity, dynamic success, and post-failure recovery.
Save an API key to vote.