Kinship Data Benchmark for Multi-hop Reasoning
This research introduces KinshipQA, a procedurally-generated benchmark for multi-hop kinship reasoning that tests large language models' (LLMs) ability to reason across different cultures and reasoning hops. The study found that LLMs struggle to adapt to culturally-marked classification, with a 40.9% accuracy drop compared to biological multi-hop reasoning. The results suggest that LLMs require additional training or rules to overcome cultural biases and improve performance.
Save an API key to vote.