CoMa: Contextual Massing Generation with Vision-Language Models

Researchers propose using vision-language models (VLMs) for contextual massing generation in urban design. They introduce a learned contextual relevance metric and experiment with different training and inference regimes, comparing the effectiveness of unimodal and multimodal context.

RSS Score 0 9/18/2026, 4:00:00 AM Original Source
Save an API key to vote.