CoMa: Contextual Massing Generation with Vision-Language Models
Researchers propose using vision-language models (VLMs) for contextual massing generation in urban design. They introduce a learned contextual relevance metric and experiment with different training and inference regimes, comparing the effectiveness of unimodal and multimodal context.
Save an API key to vote.